YOUR SCENARIO
How would you approach this?
A support-drafting feature must work within a smaller operating budget. Management proposes replacing the model for every request. How would you compare a full replacement with a routing strategy?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- Which request types create most spend, and which mistakes are expensive for operators?
- Does the budget include retries, routing calls, retrieval, and manual rework?
One defensible approach
- 01
Establish cost per useful outcome
Measure request mix, input and output tokens, cache use, failed calls, and accepted drafts. Use current provider prices with the relevant conditions. Avoid inferring production savings from headline token prices alone.
- 02
Compare policies on fixed cases
Evaluate the existing model, a cheaper candidate, and a simple routing rule on the same held-out tasks. Slice results by complexity, language, and exception handling. Include the router's cost and mistakes, not just the selected model's answer.
- 03
Release with a quality boundary
Define which failures trigger escalation and test whether the validator catches them. Introduce the policy gradually, monitor accepted outcomes and total spend, and retain a reversible configuration. Do not treat model self-confidence as calibrated correctness.
Explain the tradeoff
Routing can preserve difficult-case quality, but an inaccurate router and repeated escalation may erase the savings. A simpler model change may be better if it passes the actual workload.
Common mistakes
- Quoting a savings percentage before measuring the traffic mix.
- Evaluating only the easy cases selected for the cheaper model.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- What if the fallback rate doubles after launch?
- How would you compare a shorter prompt before changing models?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
Anthropic: Building effective agents ↗Anthropic: Demystifying evals for AI agents ↗