YOUR SCENARIO
How would you approach this?
A knowledge assistant answers the ten demo questions well. The customer wants to invite the whole company tomorrow. What evaluation would you run and how would you make the release decision?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- Which user groups and task types will the release serve, including unsupported requests?
- What failures are unacceptable and who owns the acceptance decision?
One defensible approach
- 01
Build a representative test set
Collect authorized examples from intended workflows. Include common requests, rare exceptions, unanswerable questions, conflicting sources, and access boundaries. Keep development cases separate from the held-out release check.
- 02
Use explicit grading rules
Define required evidence, acceptable answers, critical failures, and appropriate abstention with domain reviewers. Combine deterministic checks, human review, and calibrated automated graders where useful. Record model, prompt, data, and tool versions.
- 03
Make a bounded release decision
Report results by task slice alongside sample sizes, latency, and cost. For sparse evidence, propose a restricted rollout with monitoring and a fallback. Agree on stop conditions before expanding access.
Explain the tradeoff
More testing delays availability but can expose hidden workload gaps. A smaller supported release is an option when the evidence is insufficient for the full audience.
Common mistakes
- Reusing the rehearsed demo set as the only acceptance test.
- Averaging a critical permission failure into a high overall score.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- How do you add production failures without overfitting the test set?
- What changes if the system can perform actions?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
Anthropic: Demystifying evals for AI agents ↗Google SRE: Canarying releases ↗