YOUR SCENARIO
How would you approach this?
A repair assistant handles both “E-417 on ZX-8” and “the pump rattles after startup.” Vector-only retrieval handles some paraphrases but misses exact code matches. How would you improve and evaluate retrieval?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 10 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- Do extraction and indexing preserve punctuation, model numbers, and code-to-product relationships?
- How often do users provide a code, a description, or both?
One defensible approach
- 01
Build query slices
Create labeled examples for exact identifiers, paraphrases, and mixed queries. Inspect failures caused by broken text extraction or normalization before adding another retrieval layer.
- 02
Combine eligible candidates
Compare lexical and dense retrieval, then combine their ranked lists with an explicit method such as reciprocal rank fusion. Apply the same access scope to both paths. Deduplicate by stable passage IDs before assembling evidence.
- 03
Test optional reranking
Evaluate a reranker over a bounded candidate set only if the combined ranking leaves useful evidence buried. Track per-slice coverage, downstream answers, and latency; keep a simple baseline to show which stage earns its cost.
Explain the tradeoff
Hybrid retrieval introduces another index and tuning surface. It is useful only if the measured coverage gains justify its operating and query costs.
Common mistakes
- Adding raw lexical and vector scores without considering their scales.
- Expecting a reranker to recover documents absent from all candidates.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- What happens if the two result lists use different chunk IDs?
- Would you use identical settings for all query slices?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
Elasticsearch: Reciprocal rank fusion ↗Qdrant: Filtering ↗