Prepare for AI engineer and Forward Deployed Engineer (FDE) interviews with free topic questions, matched source answers, foundation practice, and customer scenarios. All questions and answers are in English, with original source numbers preserved.
An internal assistant has acceptable average latency, but some users wait more than ten seconds. It performs retrieval, reranking, generation, and optional tool calls. How do you find and improve the slow path?
Reveal an answer outline
Clarify first
Does the target describe time to first useful output or completed answers?
Measure the complete path. Record queue time and stage-level latency within a request trace. Separate first-token time from completion and include retries. Compare tail percentiles by task type and load instead of relying on a single overall average.
Test the leading bottleneck. Inspect the slow traces before changing infrastructure. If reranking dominates, test fewer candidates; if generation dominates, inspect context and output size. If the queue dominates, review concurrency and admission control. Measure the quality consequence of each change.
Protect the deadline. Give expensive stages bounded time and define acceptable degradation for optional work. Use streaming when useful to users, but still enforce a completion budget. Load-test the intended mix, including cold paths and dependency delays.
The tradeoff: Skipping a stage can improve speed while damaging a particular class of answers. Compare latency and quality for the affected slice rather than declaring success from the aggregate.
Your primary model provider starts timing out during peak traffic. A backup model is configured, but it uses different tool-call behavior and has never handled live requests. What would a credible recovery design include?
Reveal an answer outline
Clarify first
Which operations can degrade to read-only or human handling?
Bound pressure on the dependency. Set per-attempt and whole-request deadlines. Retry temporary failures with bounded backoff and jitter where safe, and stop sending new work to a repeatedly failing route. Control queues so retries do not amplify the outage.
Treat fallback as a tested product path. Evaluate the secondary model against the same acceptance suite and tool contract. Check which tasks it can safely support. For unsupported work, provide a clear degraded mode or human handoff rather than silently lowering the service's guarantees.
Exercise recovery and return. Simulate timeouts, rate limits, partial responses, and fallback exhaustion. Track routing decisions and user outcomes. Define how traffic returns to the primary gradually, and reconcile uncertain external actions before any repeat.
The tradeoff: A backup adds integration and operating cost. For some tasks, a reliable human or read-only fallback offers more value than an untested second model.
A shared RAG service stores documents for several customers. The model supplies a tenant_id argument when requesting a search, and a shared response cache keys only on the question text. Review the design.
Reveal an answer outline
Clarify first
Where does the trusted user and tenant identity originate?
Move authority outside the model. Derive identity from verified server context and enforce resource authorization at every data boundary. The model's requested tenant is not trusted authority. Use supported isolation or mandatory filters and ensure every retrieval path applies them.
Scope derived artifacts. Scope response caches to all correctness and access dimensions, including tenant, authorization state, and relevant data version, or avoid caching protected answers. Recheck citation access. Keep logs and traces from becoming a second unprotected data store.
Test with conflicting fixtures. Create two tenants with the same question but different private answers. Test direct access, semantic search, cache hits, revoked access, and malicious tool arguments. Confirm unauthorized records never enter the model context.
The tradeoff: Physical isolation can simplify some boundaries at greater operating cost. Shared storage requires consistently enforced policy across every access path; a prompt instruction cannot substitute for it.
The assistant returns HTTP 200 as usual, yet users report worse answers after yesterday's release. Several things changed: the prompt, an ingestion job, and the model alias. How do you investigate and recover?
Reveal an answer outline
Clarify first
Which task slices worsened, and what examples demonstrate the change?
Establish the impact. Collect redacted failed cases and compare task outcomes, escalation, latency, and cost with the previous period. Distinguish user-mix changes from a regression on comparable inputs. Preserve the relevant versions and traces.
Isolate likely changes. Replay a representative case set while changing one suspected component at a time where reproducible. Compare extracted documents and retrieved evidence before blaming generation. Use fixed model versions when available and record the limits of replaying a moving alias.
Recover with evidence. Use the known-good configuration if it remains safe and compatible with current data. Validate the recovery on affected tasks and monitor the rollout. Add the failures to a maintained regression set and prevent multiple untracked changes from hiding the next cause.
The tradeoff: Rollback can restore service quickly but may also restore an obsolete policy or incompatible index. Check the whole version combination rather than reversing only the application commit.