YOUR SCENARIO
How would you approach this?
A scheduling agent says it booked a meeting, but the test calendar contains no event. Another run creates the event twice. How should an evaluation score the outcome and the steps taken?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- What exact calendar state and participant constraints define success?
- Can the test environment be reset and external effects inspected independently?
One defensible approach
- 01
Verify observable state
Grade whether one intended event exists with the correct participants, time, and authorization. Check for duplicate or unwanted events. A final natural-language success message is evidence of a claim, not proof that the calendar changed.
- 02
Inspect the execution trace
Examine selected tools, arguments, results, retries, and approval handling to explain failures. Allow different valid paths unless the task requires a specific sequence. Measure wasted steps without penalizing reasonable recovery automatically.
- 03
Run isolated trials
Reset fixtures between trials, capture dependencies and versions, and repeat tasks to observe variability. Report task success and harmful side effects separately. Add the missing-event and duplicate-event cases to the regression suite.
Explain the tradeoff
A strict expected trace is easy to compare but can reject a legitimate alternative plan. Outcome checks need to capture side effects, while trace checks should explain or enforce meaningful constraints.
Common mistakes
- Letting the agent grade its own completion statement.
- Reusing a dirty calendar fixture across trials.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- How would you test a timeout after the event was created?
- What if the task has several acceptable meeting times?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
Anthropic: Demystifying evals for AI agents ↗AWS Builders’ Library: Making retries safe with idempotent APIs ↗