← Production & reliability

FREE AI FDE PRACTICE / PRODUCTION & RELIABILITY

Keep a useful service during a provider outage

Advanced 10-min practiceEditorial review: 2026-10-05

YOUR SCENARIO

How would you approach this?

Your primary model provider starts timing out during peak traffic. A backup model is configured, but it uses different tool-call behavior and has never handled live requests. What would a credible recovery design include?

This is an illustrative practice scenario. State any additional assumptions in your answer.

Make your case first.

Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.

Your notes are not submitted or saved. Keep a copy before leaving this page.

Reveal reference approach Clarifying questions, decisions, and tradeoffs

Clarify before designing.

  • Which operations can degrade to read-only or human handling?
  • Has the backup been tested for task quality, schemas, data handling, capacity, and cost?

One defensible approach

  1. 01

    Bound pressure on the dependency

    Set per-attempt and whole-request deadlines. Retry temporary failures with bounded backoff and jitter where safe, and stop sending new work to a repeatedly failing route. Control queues so retries do not amplify the outage.

  2. 02

    Treat fallback as a tested product path

    Evaluate the secondary model against the same acceptance suite and tool contract. Check which tasks it can safely support. For unsupported work, provide a clear degraded mode or human handoff rather than silently lowering the service's guarantees.

  3. 03

    Exercise recovery and return

    Simulate timeouts, rate limits, partial responses, and fallback exhaustion. Track routing decisions and user outcomes. Define how traffic returns to the primary gradually, and reconcile uncertain external actions before any repeat.

Explain the tradeoff

A backup adds integration and operating cost. For some tasks, a reliable human or read-only fallback offers more value than an untested second model.

Common mistakes

  • Assuming an API-compatible endpoint has equivalent behavior.
  • Retrying every request at every layer without a shared deadline.

KEEP THE CONVERSATION GOING

Try the follow-ups.

  1. What if both providers share the same upstream dependency?
  2. How do you prevent failover from doubling a write?

Review your own answer.

Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.

Check the underlying concepts.

The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.

AWS Builders’ Library: Making retries safe with idempotent APIs ↗Google SRE: Implementing SLOs ↗