YOUR SCENARIO
How would you approach this?
The assistant returns HTTP 200 as usual, yet users report worse answers after yesterday's release. Several things changed: the prompt, an ingestion job, and the model alias. How do you investigate and recover?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- Which task slices worsened, and what examples demonstrate the change?
- Which prompt, model, source, index, and tool versions were actually used in those runs?
One defensible approach
- 01
Establish the impact
Collect redacted failed cases and compare task outcomes, escalation, latency, and cost with the previous period. Distinguish user-mix changes from a regression on comparable inputs. Preserve the relevant versions and traces.
- 02
Isolate likely changes
Replay a representative case set while changing one suspected component at a time where reproducible. Compare extracted documents and retrieved evidence before blaming generation. Use fixed model versions when available and record the limits of replaying a moving alias.
- 03
Recover with evidence
Use the known-good configuration if it remains safe and compatible with current data. Validate the recovery on affected tasks and monitor the rollout. Add the failures to a maintained regression set and prevent multiple untracked changes from hiding the next cause.
Explain the tradeoff
Rollback can restore service quickly but may also restore an obsolete policy or incompatible index. Check the whole version combination rather than reversing only the application commit.
Common mistakes
- Treating HTTP availability as a quality metric.
- Claiming a root cause from timing alone when several components changed.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- What if the old provider model is no longer available?
- How would you change the next release process?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
Google SRE: Canarying releases ↗Anthropic: Demystifying evals for AI agents ↗