YOUR SCENARIO
How would you approach this?
A maintenance handbook mixes paragraphs, equipment tables, and footnotes. Fixed-size chunks separate a replacement interval from its equipment heading. How would you choose and test a better chunking strategy?
This is an illustrative practice scenario. State any additional assumptions in your answer.
Make your case first.
Clarify the goal, identify the biggest uncertainty, outline an approach, and explain how you would test it. Spend about 8 minutes before opening the reference.
Your notes are not submitted or saved. Keep a copy before leaving this page.
Reveal reference approach Clarifying questions, decisions, and tradeoffs
Clarify before designing.
- Which queries require multiple cells, a heading, or a footnote to be interpreted correctly?
- What structure survives the document extraction stage?
One defensible approach
- 01
Define an evidence unit
Identify complete units for the target queries: a procedure with prerequisites, or a table row with headers and applicable footnotes. Preserve source locations and document versions so a reviewer can inspect the original context.
- 02
Compare concrete candidates
Try structure-aware boundaries and bounded parent-context expansion against the current splitter. Avoid treating one token length or overlap percentage as universally correct. Keep extraction, embeddings, and evaluation queries fixed during the first comparison.
- 03
Evaluate downstream use
Label whether each retrieved result carries enough information to answer. Then inspect generated answers, context size, and latency. Include long tables and exception-heavy procedures in a held-out set before selecting the new default.
Explain the tradeoff
Larger evidence units retain meaning but may bring irrelevant rows. Smaller searchable units plus controlled surrounding context are another option to test.
Common mistakes
- Tuning chunk length while the parser has already lost table headers.
- Reporting retrieval similarity as answer correctness.
KEEP THE CONVERSATION GOING
Try the follow-ups.
- How would you handle a table spanning three pages?
- When would overlapping chunks make the result worse?
Review your own answer.
Tick the points you covered. This is a reflection checklist, not an automated score or a hiring prediction.
Check the underlying concepts.
The scenario and reference approach were written for SaveMyToken. These sources support the technical concepts; they do not report this question being asked by an employer.
LangChain: Retrieval ↗Anthropic: Effective context engineering for AI agents ↗