WHAT YOU WILL BUILD
A repeatable check for evidence lost at chunk boundaries.
A useful chunk contains the evidence needed to answer a question. In this example, the 30-day return window is incomplete without the restriction on opened items. Inspecting chunk length alone would miss that failure.
Before you start
Use a text editor and Node.js 22 or newer ↗. Check your version with node --version. Save the downloaded file in an empty folder, then open a terminal in that folder.
All data is included. No packages, account, or API key are required. This lab tests text boundaries with exact string checks. It does not run embeddings, retrieval, or an LLM.
Download the lab (.mjs)FOLLOW ALONG
Work through the example.
- 01
Read the two policy paragraphs
Find the returns window and its exception in the sample code. Write down the expected answer to: Can I return an opened item after 10 days? It is not eligible under this example policy.
- 02
Run the 40-character baseline
Download the lab below and run its command. The fixed-size split separates the evidence; its complete-rule check prints false.
- 03
Inspect the paragraph split
The second splitter uses blank lines. Print structured to see the full returns paragraph. Its check prints true because both required statements survive in one chunk.
- 04
Test the boundary again
Change width to 80 and update the baseline assertion to your predicted result. Then add a long paragraph. A production splitter also needs a size limit and a fallback for oversized sections; preserving paragraphs alone is not enough.
THE COMPLETE LAB
Run it locally.
Run this command from the folder containing your downloaded file:
node rag-chunking.mjsView or copy the complete JavaScript
// SaveMyToken local lab. Run with Node.js 22 or newer.
import assert from "node:assert/strict";
const paragraphs = [
"Returns: Unused items can be returned within 30 days. Opened items are excluded.",
"Delivery: Standard shipping takes 3 to 5 working days.",
];
const text = paragraphs.join("\n\n");
const width = 40;
const fixed = Array.from({ length: Math.ceil(text.length / width) },
(_, i) => text.slice(i * width, (i + 1) * width));
const structured = text.split(/\n\n/);
const containsRule = chunk => chunk.includes("Unused items") && chunk.includes("30 days") && chunk.includes("Opened items are excluded");
console.log("Fixed chunks preserve rule:", fixed.some(containsRule));
console.log("Paragraph chunks preserve rule:", structured.some(containsRule));
assert.equal(fixed.some(containsRule), false);
assert.equal(structured.some(containsRule), true);
Expected output for the unchanged example
Fixed chunks preserve rule: false
Paragraph chunks preserve rule: trueThe assertions also check the baseline behavior. After changing an input, predict the result and update the relevant assertion.
NOW CHANGE ONE THING
Make the example your own.
Move the exception into a separate paragraph. Both simple strategies can now lose the relationship. Add a parent section or merge the two related paragraphs, then rerun the complete-rule check.
You are done when…
You can show the exact chunk containing both the rule and exception, and explain when a paragraph-only splitter fails.
If something goes wrong
An assertion failure after changing the sample can be expected: first inspect the chunks, then update the expected result. Do not remove the check just to make the script pass.
EXPLAIN WHAT YOU LEARNED
Interview practice
Try answering aloud before opening the reference answer. These are original learning questions, not a record of any employer's interviews.
01How should you choose chunk size?
Start with the smallest passage that keeps a representative answer complete. Test those passages against a question set and the model input limit, then compare retrieval and answer quality.
Watch for: A character limit is not a token limit, and one size does not suit every document.
02What does chunk overlap solve?
Overlap repeats boundary text so a fact split between adjacent chunks may remain available. Measure whether this improves evidence coverage enough to justify the added storage and context.
Watch for: Overlap does not guarantee that a distant exception travels with its rule.
03When is parent-child retrieval useful?
Search small child passages for a precise match, then attach a bounded parent section for context. Keep the child location so the reader can verify the specific evidence.
Watch for: Returning an entire manual as the parent can erase the context savings.
04Which metadata should a chunk keep?
Keep document ID, version, section or page, and the access scope required by the application. These fields support updates, filtering, and traceable citations.
Watch for: A sequential chunk number alone may change after the source is edited.
05How do you test a splitter?
Use questions requiring a single fact, a rule with an exception, and information across a boundary. Inspect the actual chunks before measuring retrieval and final answers.
Watch for: A readable chunk is not proof that a retriever will select it.
Sources and further reading
The explanation and local exercises were written for SaveMyToken. These references support the underlying concepts; the sample outputs describe only the supplied examples.
LangChain: recursive text splitting ↗LangChain: retrieval and RAG ↗