01 / The decision: use your document structure, then test
For FAQ policies organized into self-contained paragraphs, paragraph chunks are a reasonable starting configuration. Keep fixed windows as a baseline and evaluate the questions your users actually ask. Our small authored pilot did not establish a quality winner: both configurations passed 12 of 14 strict answer-and-source checks. Paragraph chunks used 13.0% fewer input tokens, but only 3.8% less estimated money in this run because the fixed-window arm received some cache hits.
If answers cross headings or paragraphs, inspect missing evidence before choosing smaller chunks. The useful decision is whether your retrieved passages retain the conditions needed to answer correctly. This study uses lexical retrieval on fictional policies; it is not a production RAG benchmark or a model ranking.
02 / What we held fixed
We wrote six fictional English Luma policies with 24 paragraphs and froze 14 questions before calling the model. Twelve have supporting evidence; two require an unsupported/null answer. The same documents, questions, retriever, top-two retrieval limit, model settings and answer checks apply to both arms. Each question was run once per arm, alternating which arm ran first.
Fixed windows join each document into 64-word chunks with 12-word overlap. Paragraph chunks preserve each paragraph, using the same 64-word window and 12-word overlap only when a paragraph is too long. These limits count words, not model tokens. Retrieval scores unique query-word overlap plus twice the title overlap, breaking ties by chunk ID. No embeddings or reranker were used.
Measured October 6, 2026 in Asia/Shanghai; raw timestamps use UTC and fall on October 5. We requested deepseek-flash and received the same alias. Provider documentation mapped it to DeepSeek-V4.1-Flash, but the API did not expose an immutable version. Temperature 0, thinking disabled, JSON output, maximum 160 output tokens, no retries. All requests returned usable token accounting.
03 / Results: the same quality checks, different input sizes
Complete evidence means every required gold quote appears wholly inside a retrieved chunk with the correct document ID. Strict pass additionally requires the expected normalized answer, valid document IDs, all required IDs and a complete JSON response. Unsupported answers must be null with an empty source list. These checks do not replace a general semantic citation audit.
Costs below multiply observed cache-hit, cache-miss and output tokens by both published tariffs: off-peak and peak. They are estimates for the entire listed group, not invoices or per-request prices. The October 6 snapshot is USD 0.006 / 0.30 / 1.20 per million tokens at peak, and half those rates off-peak. We do not infer holiday eligibility or actual account billing. Totals include failed answer checks; values are rounded for display.
| Metric | Fixed windows | Paragraph chunks |
|---|---|---|
| Strict answer + source pass | 12/14 | 12/14 |
| Complete evidence (answerable questions) | 12/12 | 12/12 |
| Correct unsupported/null answers | 2/2 | 2/2 |
| Total input tokens | 4355 | 3789 |
| Cached input tokens | 384 | 0 |
| Output tokens | 207 | 209 |
| Estimated total USD (off-peak–peak) | $0.00072100–$0.00144200 | $0.00069375–$0.00138750 |
| Median observed request duration | 543 ms | 620.5 ms |
04 / The failures were source identifiers
Both arms failed q01 and q02 despite giving the expected answer values. For q01 the model returned 30 days but cited Luma returns policy rather than the document ID returns. The q02 failures likewise used a title or an ID/title combination. We retained all four failures under the original grading contract. Finding the evidence and producing a usable citation are separate requirements.
A next iteration should distinguish the document ID and title more explicitly in the prompt, then rerun every question under a new experiment version. Changing the checker after seeing the answers would make the published comparison harder to interpret.
Expected: {"answer":"30","sourceIds":["returns"]}
Observed in both q01 arms: {"answer":"30","sourceIds":["Luma returns policy"]}05 / Reproduce and adapt the comparison
Download and extract the kit. Node.js 22 or newer is sufficient; no packages are required. The command below runs retrieval locally and makes no API call. For a measured answer comparison, follow the README and run with --live, your own private environment file, and a new --out path. Answer keys and evidence labels are used locally for scoring, never sent to the API.
For your own evaluation, freeze representative documents and questions first. Include exceptions, facts spanning boundaries, near-matching policies and unsupported questions. Record evidence coverage, strict passes, full input/output/cache usage and failed attempts before choosing a configuration.
node run.mjs --experiment faq --out local-retrieval.json06 / What this pilot cannot establish
Fourteen authored questions are too few to generalize across document collections. Our word-based lexical retriever is intentionally simple. We tested one model alias and one size/overlap setting, without confidence intervals. The FAQ and cache experiments ran concurrently, with up to two requests in flight, so request durations include network and provider-load effects. No speed advantage is established.
The two unanswerable questions demonstrate two successful abstentions, not a broad hallucination rate. Reproduce the procedure on a larger held-out dataset before applying a result to customers. A configuration that preserves these short paragraphs may fail on tables, long manuals or cross-document questions.