YOUR SCENARIO
How would you approach this?
A research agent keeps reformulating the same failed search and spends its entire daily allowance on one task. It stays under a per-call timeout, so the existing timeout check never stops it. What controls are missing?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 8 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- What counts as a completed task, and which failures should terminate it?
- Are limits shared by nested agents, retries, and parallel calls?
One defensible approach
- 01
Budget the whole run
Set limits for total calls, elapsed time, and spend at the controller boundary. Include child work and retries in the same accounting scope. Reserve estimated cost before concurrent work so multiple branches cannot all consume the remaining budget.
- 02
Require meaningful progress
Track action identity, results, and changes in evidence. Detect repeated unsuccessful actions while allowing legitimate pagination or polling. Require a new hypothesis, new input, or escalation when the same approach no longer improves the task state.
- 03
Stop with usable state
Cancel pending work where supported, record completed actions, and distinguish budget exhaustion from success. Return bounded findings and remaining gaps. Test hung tools, concurrent branches, and misleading success messages.
Explain the tradeoff
Tight limits can stop valid long investigations. Make the budget visible and adjustable for the task while retaining a hard ceiling and an explicit stop reason.
Common mistakes
- Relying only on the model to decide when enough money was spent.
- Applying separate full budgets to every child agent.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- Can cancellation guarantee that remote work stops immediately?
- How do you distinguish a useful retry from a loop?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
Anthropic: Building effective agents ↗Anthropic: Demystifying evals for AI agents ↗