# SaveMyToken FAQ and prompt-cache pilot — 2026-10-06

Node.js 22 or newer; no npm packages required. Extract `experiment-kit.zip` into a directory. Keep `run.mjs`, `core.mjs`, and `faq-dataset.json` together.

## Inspect without an API call

```sh
node run.mjs --experiment faq --out local-retrieval.json
node run.mjs --experiment cache --out cache-plan.json
```

FAQ defaults to local retrieval only. Cache defaults to a request count/plan summary. Neither mode calls a model or reports measured answer quality/cost. Output paths must not already exist. Reading the articles or downloading this packet makes no model requests.

## Repeat the measured procedure

Create a private environment file outside your public/site directory containing your own `DEEPSEEK_API_KEY` and optional `DEEPSEEK_MODEL=deepseek-flash`. Never commit or share it. Then run:

```sh
node --env-file=/absolute/path/to/private.env run.mjs --experiment faq --live --out my-faq.json
node --env-file=/absolute/path/to/private.env run.mjs --experiment cache --live --out my-cache.json
```

Live mode makes billable requests to `https://api.deepseek.com/chat/completions`. FAQ plans 28; cache plans 24. Output is capped at 160 tokens/request, temperature 0, thinking disabled, JSON output. The runner uses no retries, saves progress after every response, and stops on a transport/API error or missing usage. Unknown usage remains unknown, including potential charges for failed requests. Check current provider prices before repeating: the script deliberately retains the October 6 snapshot.

A conservative preflight estimate counts UTF-8 input bytes plus overhead as tokens, assumes no cache hits, and includes the output limit. Plans over USD 0.25 per experiment are refused. This is not a provider-enforced account spending cap.

## Frozen dataset and methods

Six authored fictional English policies, 24 paragraphs, 14 questions: 12 answerable and two unsupported. The original dataset is dedicated to the public domain under CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/ . It is not a real store policy or representative customer sample.

FAQ: join each document into 64-word windows with 12-word overlap, or preserve paragraphs with the same window/overlap fallback for long paragraphs. These are words, not model tokens. Retrieve top two positive lexical matches: unique query-word overlap plus twice title overlap; ties by numeric chunk ID. Alternate which arm runs first per question. No embeddings, reranker, model changes or post-run answer-key edits.

Cache: send all six documents, using q01–q06 in two blocks (second block reversed), alternating the leading arm. Stable system prompt; document-before-question vs question-before-document in the user message. A run/block/arm marker precedes each prompt, isolating each trial prefix beyond the common system message. Each arm has two initial and ten later requests. Calls within a process are sequential with 1.5 seconds between them. The two published experiment processes ran concurrently, up to two requests in flight. Run separately for a less confounded latency study.

Strict pass: complete JSON, accepted normalized answer, valid document IDs, all required IDs, and empty IDs for an unsupported/null answer. Evidence coverage separately checks whether every required verbatim quote occurs wholly in a retrieved chunk. Neither measure is an unrestricted semantic citation judge. Answer keys and evidence fields are never sent to the API.

## Read the results

- `faq-results.json`, `cache-results.json`: every sent message, raw answer, usage, elapsed request time, finish reason and verdict, plus summaries. No credentials or response/account IDs.
- `faq-results.csv`, `cache-results.csv`: one row per request for analysis; JSON is the complete record.
- `manifest.json`: SHA-256 for evidence and reproduction files, measured source checkpoint, provenance.

Run time: 2026-10-06 Asia/Shanghai (UTC timestamps in JSON fall on October 5). Requested and returned alias `deepseek-flash`; documentation mapped it to DeepSeek-V4.1-Flash when checked, but the API exposed no immutable model version.

FAQ: both arms pass 12/14, cover evidence for 12/12 answerable questions, and abstain on 2/2 unsupported questions. All four strict failures gave the right value but invalid source IDs. Paragraph input was smaller, with no demonstrated quality win.

Cache: document-first 12/12 vs question-first 11/12; later requests both 10/10. Document-first later cache hits were 12,800/15,513 tokens; question-first had zero. Initial requests had zero hits in both arms. The one incorrect answer is retained. Document-first was slower by median request duration; this pilot does not establish a speed improvement.

Cost = (cache-hit tokens × hit rate + cache-miss tokens × miss rate + output tokens × output rate) / 1,000,000. Peak rates USD 0.006 / 0.30 / 1.20; off-peak USD 0.003 / 0.15 / 0.60. Both totals are estimates, not invoices. No holiday or account eligibility is inferred. Source: https://api-docs.deepseek.com/quick_start/pricing/ . Cache behavior: https://api-docs.deepseek.com/guides/kv_cache . Checked October 6, 2026.

Small authored pilot, one model alias, short answers, no confidence intervals. Different documents, boundaries, retrieval, versions, cache retention, timing and provider load can change results. Reproduction means rerunning the same procedure, not guaranteeing identical answers or cache hits.
