A research agent is only useful if you can trust what it says and bound what it spends. Here a planner splits a question into sub-questions, researchers search and read pages with tools that fail on purpose, a supervisor sends unanswered ones back once, and a writer cites every claim and names what it could not find. Budgets are enforced in code, and the state is saved after every step. Below: real runs on a local copy of PyPI pages, at three tool failure rates.
Each step is one agent's turn. Sub-questions turn green when answered with a citation; a failed tool call leaves them for the supervisor, which sends the retryable ones back once. The writer reports what it could not establish instead of guessing.
The report
More failures, fewer answers, never a wrong one
All 100 questions (600 sub-questions) at each tool failure rate. Unanswered sub-questions are named in the report; claims that were made all matched the page they cite.