Prepare for AI engineer and Forward Deployed Engineer (FDE) interviews with free topic questions, matched source answers, foundation practice, and customer scenarios. All questions and answers are in English, with original source numbers preserved.
A knowledge assistant answers the ten demo questions well. The customer wants to invite the whole company tomorrow. What evaluation would you run and how would you make the release decision?
Reveal an answer outline
Clarify first
Which user groups and task types will the release serve, including unsupported requests?
Build a representative test set. Collect authorized examples from intended workflows. Include common requests, rare exceptions, unanswerable questions, conflicting sources, and access boundaries. Keep development cases separate from the held-out release check.
Use explicit grading rules. Define required evidence, acceptable answers, critical failures, and appropriate abstention with domain reviewers. Combine deterministic checks, human review, and calibrated automated graders where useful. Record model, prompt, data, and tool versions.
Make a bounded release decision. Report results by task slice alongside sample sizes, latency, and cost. For sparse evidence, propose a restricted rollout with monitoring and a fallback. Agree on stop conditions before expanding access.
The tradeoff: More testing delays availability but can expose hidden workload gaps. A smaller supported release is an option when the evidence is insufficient for the full audience.
Your automatic evaluator prefers a verbose answer that adds an unsupported claim over a concise correct answer. Its scores disagree with support reviewers. How would you repair the evaluation?
Reveal an answer outline
Clarify first
What dimensions is the judge supposed to measure, and does it have the actual source evidence?
Separate the criteria. Create a rubric for support, completeness, relevance, and style instead of one vague quality score. Add concrete positive and negative examples, including fluent answers with false claims. Decide how critical errors affect acceptance.
Measure judge behavior. Use an independently reviewed set. For pairwise grading, swap answer order and check consistency. Examine false acceptances, false rejections, and disagreements by category. Mask irrelevant model identity and test whether length alone changes a decision.
Version and audit the grader. Record the judge model, prompt, rubric, and evidence version. Recalibrate after changes and route important disagreements to people. Treat a model-generated score as a measurement needing validation rather than ground truth.
The tradeoff: Human review is costly and can disagree too. Spend it on rubric design, representative calibration, and consequential disagreements rather than every easy case.
A scheduling agent says it booked a meeting, but the test calendar contains no event. Another run creates the event twice. How should an evaluation score the outcome and the steps taken?
Reveal an answer outline
Clarify first
What exact calendar state and participant constraints define success?
Verify observable state. Grade whether one intended event exists with the correct participants, time, and authorization. Check for duplicate or unwanted events. A final natural-language success message is evidence of a claim, not proof that the calendar changed.
Inspect the execution trace. Examine selected tools, arguments, results, retries, and approval handling to explain failures. Allow different valid paths unless the task requires a specific sequence. Measure wasted steps without penalizing reasonable recovery automatically.
Run isolated trials. Reset fixtures between trials, capture dependencies and versions, and repeat tasks to observe variability. Report task success and harmful side effects separately. Add the missing-event and duplicate-event cases to the regression suite.
The tradeoff: A strict expected trace is easy to compare but can reject a legitimate alternative plan. Outcome checks need to capture side effects, while trace checks should explain or enforce meaningful constraints.
An extraction service always returns parseable JSON, yet it omits required fields and occasionally proposes a negative quantity. The team says JSON mode guarantees correctness. How would you redesign validation?
Reveal an answer outline
Clarify first
Does the selected API mode guarantee JSON syntax or adherence to a supported schema?
State the guarantees accurately. Distinguish JSON syntax from schema adherence. Where supported, use a strict structured-output schema for required fields, types, and permitted values. Verify the provider's supported subset rather than assuming every JSON Schema keyword is enforced.
Validate at the application boundary. Parse and validate responses before downstream use. Check quantity ranges, record existence, units, and relevant business rules with authoritative data. A schema-valid customer ID still does not establish authorization to act on that customer.
Handle non-success explicitly. Recognize refusals, incomplete responses, and transport errors. Use bounded correction or review for recoverable issues, and reject unsafe output before side effects. Test malformed, incomplete, schema-valid-but-wrong, and adversarial cases.
The tradeoff: Stricter constraints reduce structural variation but cannot determine factual truth. Complex business checks often belong outside generation where they can be tested deterministically.