← Production & reliability

FREE AI FDE PRACTICE / PRODUCTION & RELIABILITY

The average is fast. Why are users waiting?

Intermediate 10-min practiceEditorial review: 2026-10-05

YOUR SCENARIO

How would you approach this?

An internal assistant has acceptable average latency, but some users wait more than ten seconds. It performs retrieval, reranking, generation, and optional tool calls. How do you find and improve the slow path?

This is an illustrative practice scenario. State any additional assumptions in your answer.

Make your case first.

Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.

Your notes are not submitted or saved. Keep a copy before leaving this page.

Reveal reference approach Clarifying questions, decisions, and tradeoffs

Clarify before designing.

  • Does the target describe time to first useful output or completed answers?
  • Which users, payload sizes, and concurrency levels appear in the slow tail?

One defensible approach

  1. 01

    Measure the complete path

    Record queue time and stage-level latency within a request trace. Separate first-token time from completion and include retries. Compare tail percentiles by task type and load instead of relying on a single overall average.

  2. 02

    Test the leading bottleneck

    Inspect the slow traces before changing infrastructure. If reranking dominates, test fewer candidates; if generation dominates, inspect context and output size. If the queue dominates, review concurrency and admission control. Measure the quality consequence of each change.

  3. 03

    Protect the deadline

    Give expensive stages bounded time and define acceptable degradation for optional work. Use streaming when useful to users, but still enforce a completion budget. Load-test the intended mix, including cold paths and dependency delays.

Explain the tradeoff

Skipping a stage can improve speed while damaging a particular class of answers. Compare latency and quality for the affected slice rather than declaring success from the aggregate.

Common mistakes

  • Calling a fast first token proof of a fast completed task.
  • Optimizing the model when most time is spent waiting in a queue.

KEEP THE CONVERSATION GOING

Try the follow-ups.

  1. How do you choose a latency objective with the customer?
  2. What happens when the optional tool times out?

Review your own answer.

Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.

Check the underlying concepts.

The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.

Google SRE: Implementing SLOs ↗LangChain: Retrieval ↗