← Context & cost

FREE AI FDE PRACTICE / CONTEXT & COST

Reduce model spend without hiding quality losses

Advanced 10-min practiceEditorial review: 2026-10-05

YOUR SCENARIO

How would you approach this?

A support-drafting feature must work within a smaller operating budget. Management proposes replacing the model for every request. How would you compare a full replacement with a routing strategy?

This is an illustrative practice scenario. State any additional assumptions in your answer.

Make your case first.

Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.

Your notes are not submitted or saved. Keep a copy before leaving this page.

Reveal reference approach Clarifying questions, decisions, and tradeoffs

Clarify before designing.

  • Which request types create most spend, and which mistakes are expensive for operators?
  • Does the budget include retries, routing calls, retrieval, and manual rework?

One defensible approach

  1. 01

    Establish cost per useful outcome

    Measure request mix, input and output tokens, cache use, failed calls, and accepted drafts. Use current provider prices with the relevant conditions. Avoid inferring production savings from headline token prices alone.

  2. 02

    Compare policies on fixed cases

    Evaluate the existing model, a cheaper candidate, and a simple routing rule on the same held-out tasks. Slice results by complexity, language, and exception handling. Include the router's cost and mistakes, not just the selected model's answer.

  3. 03

    Release with a quality boundary

    Define which failures trigger escalation and test whether the validator catches them. Introduce the policy gradually, monitor accepted outcomes and total spend, and retain a reversible configuration. Do not treat model self-confidence as calibrated correctness.

Explain the tradeoff

Routing can preserve difficult-case quality, but an inaccurate router and repeated escalation may erase the savings. A simpler model change may be better if it passes the actual workload.

Common mistakes

  • Quoting a savings percentage before measuring the traffic mix.
  • Evaluating only the easy cases selected for the cheaper model.

KEEP THE CONVERSATION GOING

Try the follow-ups.

  1. What if the fallback rate doubles after launch?
  2. How would you compare a shorter prompt before changing models?

Review your own answer.

Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.

Check the underlying concepts.

The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.

Anthropic: Building effective agents ↗Anthropic: Demystifying evals for AI agents ↗