# Agent context and cache pilot — October 6, 2026

This is the public, sanitized evidence packet for [Agent context, tokens, and cache cost](/guides/agent-context-cache-cost). The fixture, exact user prompt, per-configuration usage summaries, and DeepSeek price snapshot are included. No credentials, session transcripts, local machine paths, or private configuration are included.

## Reproduce the task

1. Copy `package.json`, `src/`, `test/`, and `PROMPT.txt` into a fresh directory for each run. Use Node.js 22 or newer. Do not install packages.
2. Run `npm test` before invoking an agent: the whitespace-run test should fail. Each agent must inspect the two files, edit only `src/normalize-label.js`, run `npm test`, and report the result. The exact user prompt is in `PROMPT.txt`.
3. Use the profile and model shown in `results.json`. Point Pi and both DSH profiles directly to DeepSeek `deepseek-flash`. Claude Code uses DeepSeek's documented Anthropic-compatible endpoint and `deepseek-flash[1m]`; no Anthropic model was used. Codex uses its OpenAI account and `gpt-5.6-sol`. Hermes was requested with `deepseek-flash` but recorded `deepseek-chat` after local normalization. Provide your own credentials through your local secret store, never in the fixture.
4. Use one fresh directory per run. The DSH profiles also need separate fresh `DSH_HOME` directories. For Pi, disable sessions, context files, extensions, skills, prompt templates, and themes. For DSH, `sdk-minimal` is the official JSON-RPC SDK profile with its minimal tree; `headless` is the separate official headless profile. For Claude Code, compare the existing local configuration with the same invocation plus `--bare`. Use only `Read`, `Edit`, and `Bash(npm test)` in the Claude invocations. Codex uses `--ephemeral`, `--sandbox workspace-write`, `--skip-git-repo-check`, and explicit `--model gpt-5.6-sol`. Hermes uses the terminal toolset, safe mode, and a 12-turn cap.
5. Collect token usage from each tool's output or local session store. The `usageSource` field identifies which record was used. `requestUsage` arrays, where available, are `[uncached input, cached input, cache write input, output]` for each model request. They are usage records, not reconstructed system prompts.

All seven completed runs changed the implementation to `replace(/\s+/g, "-")` and then reported three passing tests and zero failures. Each configuration ran **once**. There were no repeated trials, confidence intervals, or hidden tests. A successful three-test patch is not a general coding-quality evaluation.

## Token definitions and estimated cost

`uncachedInputTokens` and `cachedInputTokens` are disjoint in this file; total input is their sum plus `cacheWriteInputTokens`. Pi and DSH expose those categories directly. Claude Code's JSON `input_tokens` excludes `cache_read_input_tokens`, so we add them for total input. Codex's `input_tokens` already includes `cached_input_tokens`, so uncached is the difference. Hermes exposes session-level `input_tokens` and `cache_read_tokens`; we use them as distinct categories. Different providers and clients may tokenize or classify usage differently.

Cache-hit share = cached input / total input. It is an observed share for a run, not a fixed property of an agent. The DeepSeek list-rate estimate for runs with a verified Flash alias is:

```
(uncached input × cache-miss rate
 + cached input × cache-hit rate
 + output × output rate) / 1,000,000
```

The October 6 snapshot in `results.json` contains peak and off-peak rates. Rates may change. The estimate is **not an invoice**: it does not prove the account's billed time band, discounts, or other charges. The `total_cost_usd` field printed by Claude Code was not used because its pricing metadata did not match this DeepSeek setup. We do not apply DeepSeek Flash prices to Codex's GPT model or Hermes's unresolved `deepseek-chat` alias. If a provider bills cache writes or tools separately, add those charges; the included CLI records reported zero cache-write tokens.

## What this pilot can and cannot show

The two DSH runs share version `0.2.0-rc.2`, provider, model name, fixture, and user prompt. They use different official profiles, which change context and tools. The two Claude Code runs share version, provider, model name, fixture, and invocation apart from `--bare`; the regular run used the machine's existing local configuration. It was not a clean product-default installation. Turns, tool choices, file contents returned by tools, provider cache state, and response lengths also differed. These are real workload totals for these particular runs, not isolated measurements of built-in system-prompt length.

Pi was deliberately stripped of optional context and resources; it is not its default configuration. DSH `sdk-minimal` is an SDK profile, not the Web UI's Minimal preset. Codex uses a different model and account billing path. Hermes's CLI normalized `deepseek-flash` to `deepseek-chat`, and we did not get an immutable served-model identifier. Cache states were not reset at the provider. A longer, more varied task suite with repeated runs and controlled tool/configuration baselines is needed before ranking agent products by cost per successful task.

## Sources

- DeepSeek model aliases and prices: https://api-docs.deepseek.com/quick_start/pricing/
- DeepSeek's Claude Code integration: https://api-docs.deepseek.com/quick_start/agent_integrations/claude_code/
- DSH `sdk-minimal` profile: https://github.com/deepseek-ai/deepseek-harness/blob/master/packages/bundle/sdk-minimal/README.md
- Pi's context and resource assembly: https://pi.dev/docs/latest/how-pi-works
- Anthropic cache usage semantics: https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- OpenAI cache accounting: https://developers.openai.com/api/docs/guides/prompt-caching
