Prepare for AI engineer and Forward Deployed Engineer (FDE) interviews with free topic questions, matched source answers, foundation practice, and customer scenarios. All questions and answers are in English, with original source numbers preserved.
A travel-planning assistant exceeds its context budget after many revisions. A summary reduces tokens but forgets that the traveler cannot use stairs. How would you redesign the retained context?
Reveal an answer outline
Clarify first
Which facts are binding constraints, which were corrected, and which are historical discussion?
Separate state from narration. Store explicit current constraints, decisions, and unresolved questions in a scoped task record. Retain provenance and corrections. Use a conversational summary for background, not as the sole store of requirements that determine acceptable results.
Build a bounded prompt. Keep the current request, binding constraints, recent relevant turns, and targeted evidence. Budget tool results and output space as well as messages. Retrieve older detail when needed rather than keeping every previous itinerary in the active context.
Test retention under revision. Create long conversations with corrected preferences, distractors, and conflicting old plans. Check whether generated itineraries satisfy the latest constraints after compression. Compare token usage and task success on the same conversations.
The tradeoff: Structured state takes more application work and can itself be wrong. Validate updates and retain enough provenance to inspect how a constraint entered the record.
A user changes a project preference from Python to TypeScript. The assistant acknowledges the change, but a later session retrieves the old preference from vector memory. How would you design reliable corrections?
Reveal an answer outline
Clarify first
Is the preference global, specific to one project, or only for the current task?
Define a scoped record. Represent the owner, project scope, preference key, revision, source, and validity. Derive user identity from authenticated context. A retrieved sentence should not override a more authoritative correction simply because it ranks highly.
Propagate the correction. Supersede the old value in the authoritative store and update or exclude its derived retrieval records. Invalidate affected cached context. Preserve an audit trail only under an appropriate retention policy, not as another active preference candidate.
Test the complete lifecycle. Verify recall before and after correction, across sessions and projects. Test expiration, deletion, ambiguous scope, and other users' records. Ask for clarification when a correction does not identify which project it applies to.
The tradeoff: Keeping more memories can improve recall but increases conflict and maintenance work. Store information with a defined use and lifecycle.
A support-drafting feature must work within a smaller operating budget. Management proposes replacing the model for every request. How would you compare a full replacement with a routing strategy?
Reveal an answer outline
Clarify first
Which request types create most spend, and which mistakes are expensive for operators?
Establish cost per useful outcome. Measure request mix, input and output tokens, cache use, failed calls, and accepted drafts. Use current provider prices with the relevant conditions. Avoid inferring production savings from headline token prices alone.
Compare policies on fixed cases. Evaluate the existing model, a cheaper candidate, and a simple routing rule on the same held-out tasks. Slice results by complexity, language, and exception handling. Include the router's cost and mistakes, not just the selected model's answer.
Release with a quality boundary. Define which failures trigger escalation and test whether the validator catches them. Introduce the policy gradually, monitor accepted outcomes and total spend, and retain a reversible configuration. Do not treat model self-confidence as calibrated correctness.
The tradeoff: Routing can preserve difficult-case quality, but an inaccurate router and repeated escalation may erase the savings. A simpler model change may be better if it passes the actual workload.
Your app sends a long policy and tool definitions on every request. Caching was enabled, but billed usage barely changed. Each request also includes a timestamp and user-specific settings near the beginning. How do you investigate?
Reveal an answer outline
Clarify first
What does this provider cache, and what are its eligibility, lifetime, write, and read billing rules?
Inspect the request shape. Compare consecutive serialized requests. Look for changing timestamps, tool order, formatting, and user data inside the intended reusable prefix. Confirm the endpoint and model support the requested caching behavior.
Separate stable and variable content. Place eligible shared instructions and tool definitions in a stable order, with per-request content afterward where the provider's rules permit. Keep version changes deliberate. Do not move sensitive data across users merely to improve reuse.
Measure warm and cold behavior. Run controlled repeated requests within the relevant cache lifetime. Inspect cache-specific usage and compare total billed reads, writes, and uncached tokens. Report results for this workload rather than assuming a universal discount.
The tradeoff: Maintaining a stable prefix can improve reuse but must not freeze a policy that needs updating. Correctness and version invalidation remain part of the design.