← RAG & retrieval

FREE AI FDE PRACTICE / RAG & RETRIEVAL

Retrieve exact product codes and everyday language

Intermediate 10-min practiceEditorial review: 2026-10-05

YOUR SCENARIO

How would you approach this?

A repair assistant handles both “E-417 on ZX-8” and “the pump rattles after startup.” Vector-only retrieval handles some paraphrases but misses exact code matches. How would you improve and evaluate retrieval?

This is an illustrative practice scenario. State any additional assumptions in your answer.

Make your case first.

Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.

Your notes are not submitted or saved. Keep a copy before leaving this page.

Reveal reference approach Clarifying questions, decisions, and tradeoffs

Clarify before designing.

  • Do extraction and indexing preserve punctuation, model numbers, and code-to-product relationships?
  • How often do users provide a code, a description, or both?

One defensible approach

  1. 01

    Build query slices

    Create labeled examples for exact identifiers, paraphrases, and mixed queries. Inspect failures caused by broken text extraction or normalization before adding another retrieval layer.

  2. 02

    Combine eligible candidates

    Compare lexical and dense retrieval, then combine their ranked lists with an explicit method such as reciprocal rank fusion. Apply the same access scope to both paths. Deduplicate by stable passage IDs before assembling evidence.

  3. 03

    Test optional reranking

    Evaluate a reranker over a bounded candidate set only if the combined ranking leaves useful evidence buried. Track per-slice coverage, downstream answers, and latency; keep a simple baseline to show which stage earns its cost.

Explain the tradeoff

Hybrid retrieval introduces another index and tuning surface. It is useful only if the measured coverage gains justify its operating and query costs.

Common mistakes

  • Adding raw lexical and vector scores without considering their scales.
  • Expecting a reranker to recover documents absent from all candidates.

KEEP THE CONVERSATION GOING

Try the follow-ups.

  1. What happens if the two result lists use different chunk IDs?
  2. Would you use identical settings for all query slices?

Review your own answer.

Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.

Check the underlying concepts.

The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.

Elasticsearch: Reciprocal rank fusion ↗Qdrant: Filtering ↗