01 / The decision: put reusable documents before changing questions
When many questions use the same documents, start with stable instructions and documents first, then append the current question. Verify the provider's cache-hit usage instead of assuming reuse. In this pilot, document-first prompts had an estimated total cost 64.9% below question-first prompts across all 12 requests per arm. The reduction was 77.9% for the ten later requests per arm. Initial requests had no measured cache hits.
This is a measured result for a small repeated-document workload, not a guaranteed saving. The provider still generated each answer and billed output tokens. We did not observe a speed improvement: document-first requests had a higher median duration in this run.
02 / The paired setup
We used all six fictional policies from the FAQ dataset, sending the same document content and six questions in each arm. One layout puts documents before the question; the other puts the question before documents. The system instruction is identical. Two blocks reverse question order and alternate which arm runs first. Each arm therefore has two initial requests and ten later requests.
A unique run/block/arm marker precedes the user message to separate the trial prefixes. That marker differs between arms; beyond it, the intended comparison is document/question order. The common system prefix may still be cached. Requests inside each process are sequential with 1.5 seconds between them. “Initial” describes position in a block, not a guarantee about the provider's internal cache state.
Measured October 6, 2026 in Asia/Shanghai; raw timestamps use UTC and fall on October 5. We requested deepseek-flash and received the same alias. Provider documentation mapped it to DeepSeek-V4.1-Flash, but the API did not expose an immutable version. Temperature 0, thinking disabled, JSON output, maximum 160 output tokens, no retries. All requests returned usable token accounting.
Document-first: stable system → trial marker → documents → current question
Question-first: stable system → trial marker → current question → documents03 / All requests, including the initial calls
Both arms sent 18,616 input tokens and produced 180 output tokens. Document-first reused 12,800 input tokens; question-first reported none. That isolates a large accounting difference without pretending the prompts were shorter.
Costs below multiply observed cache-hit, cache-miss and output tokens by both published tariffs: off-peak and peak. They are estimates for the entire listed group, not invoices or per-request prices. The October 6 snapshot is USD 0.006 / 0.30 / 1.20 per million tokens at peak, and half those rates off-peak. We do not infer holiday eligibility or actual account billing. Totals include failed answer checks; values are rounded for display.
| Metric across 12 requests | Document-first | Question-first |
|---|---|---|
| Strict answer + source pass | 12/12 | 11/12 |
| Input tokens | 18616 | 18616 |
| Cached input tokens | 12800 | 0 |
| Input cache-hit share | 68.8% | 0.0% |
| Estimated total USD (off-peak–peak) | $0.00101880–$0.00203760 | $0.00290040–$0.00580080 |
| Median observed request duration | 711 ms | 542 ms |
04 / Initial and later requests tell different stories
The initial two requests per arm had zero cache-hit tokens and identical estimated totals. For the ten later requests, document-first reused 82.5% of input tokens and both arms passed all ten strict checks. Keeping the initial cost visible matters when estimating a workload with frequent document changes or short sessions.
| Group / metric | Document-first | Question-first |
|---|---|---|
| Initial: requests | 2 | 2 |
| Initial: cached / input tokens | 0 / 3103 | 0 / 3103 |
| Initial: estimated total USD | $0.00048345–$0.00096690 | $0.00048345–$0.00096690 |
| Later: requests | 10 | 10 |
| Later: cached / input tokens | 12800 / 15513 | 0 / 15513 |
| Later: strict pass | 10/10 | 10/10 |
| Later: estimated total USD | $0.00053535–$0.00107070 | $0.00241695–$0.00483390 |
05 / One wrong answer and no measured speed win
Question-first failed q06 in the second block's initial request: it answered yes to outdoor warranty coverage even though the supplied policy excluded outdoor use. That response and its cost remain in the results. One error in this pilot does not establish that document-first generally improves answer quality.
Document-first median duration was 711 ms versus 542 ms for question-first. The FAQ and caching processes ran concurrently, with up to two requests in flight. Small samples, network conditions and provider load prevent a clean speed inference. Lower estimated cost is the supported finding here; a latency claim would need a separate, larger controlled run.
06 / Reproduce with your own documents
Download and extract the kit; Node.js 22 or newer is sufficient. The command below produces a plan summary without calling a model. The README shows how to use --live with your own private environment file and a new output path. Live mode is billable, capped at 24 planned requests, with no retries. Viewing this page or downloading the files makes no provider calls.
For your own workload, freeze the system instructions and document order, then vary only the question. Keep timestamps, per-request identifiers and other changing metadata out of the shared prefix. Record cache-hit/miss input and output separately. Repeat after realistic idle gaps and document changes; measure answer quality and all initial costs alongside savings.
node run.mjs --experiment cache --out cache-plan.json07 / Limits and when the result may not transfer
This pilot has two blocks, six authored questions, short outputs and one provider/model alias. It does not measure retention over long gaps, different model versions, production traffic or confidence intervals. Cache hits can vary across runs, and a changing document prefix can eliminate reuse. Repeating the procedure does not guarantee the same responses or cache-hit totals.
When each question needs different documents, evaluate retrieval and smaller context alongside caching. When outputs are long, output charges can dominate. Estimate from your actual mix of initial requests, repeated requests, failures and output lengths before adopting a savings target.