01 / One tiny fix, seven agent configurations

We gave fresh copies of the same three-test Node.js fixture to Pi, DeepSeek Harness (DSH), Claude Code, Codex, and Hermes. Every run fixed a whitespace-normalization bug and finished with 3/3 tests passing. Their input accounting still ranged from 3,161 tokens for DSH sdk-minimal to 191,383 for Claude Code with this machine's existing local configuration. The spread shows why an agent's context assembly belongs in a token budget. It does not rank the agents' coding ability.

Two comparisons are especially useful. With DSH 0.2.0-rc.2 and the same DeepSeek Flash model, sdk-minimal consumed 3,161 total input tokens while headless consumed 53,470. With Claude Code 2.1.226 using DeepSeek's Anthropic-compatible endpoint, --bare consumed 9,411 total input tokens while the existing local configuration consumed 191,383. These are one-run configuration differences, not isolated measurements of system-prompt length or clean product defaults.

02 / What we ran and counted

The fixture contains one buggy JavaScript function and three tests. The prompt asked each agent to inspect the function and tests, collapse each internal whitespace run to one hyphen, run npm test, and avoid other directories or package changes. We used a fresh fixture directory for every configuration and retained all successful runs. Each configuration ran once on October 6, 2026. The full fixture, exact prompt, usage records and method are available below.

Pi used an isolated configuration with sessions, context files, extensions, skills and templates disabled. DSH's official sdk-minimal JSON-RPC profile and its headless profile used separate clean homes. Claude Code used the same DeepSeek Flash endpoint in both runs; --bare strips local customization, while the ordinary run used this machine's existing setup. Codex used GPT-5.6 Sol through a ChatGPT account. Hermes was requested with deepseek-flash, but its CLI recorded deepseek-chat after normalizing the alias. All five products actually ran; Hermes's served model version remains unresolved.

For a run, total input means uncached input + cache reads + reported cache writes. Cache-hit share means cache reads divided by that total. Pi and DSH emitted per-request usage. Claude Code and Codex emitted aggregate usage; Hermes supplied session counters. The clients do not all name or aggregate fields identically, so the public data explains each conversion.

03 / Measured tokens and illustrative list-rate cost

The cost column applies the DeepSeek Flash rates checked on October 6 to rows whose Flash alias we could verify: cache hits $0.003–$0.006, misses $0.15–$0.30, and output $0.60–$1.20 per million tokens, depending on off-peak versus peak. It is a list-rate estimate for the full run, not an invoice. We do not price Codex with a DeepSeek tariff or price Hermes's unresolved model alias. Claude Code's own total_cost_usd field used incompatible price metadata for this endpoint, so it is excluded. All seven runs passed the same three tests.

DeepSeek Flash pricing and served-model alias, checked October 6, 2026 ↗

Measured tokens and illustrative list-rate cost
Agent / configurationTotal inputCache read / shareOutputEst. USD, off-peak–peak
Pi 0.85.1 · stripped resources · Flash10,1067,680 / 76.0%422$0.000640–$0.001280
DSH 0.2.0-rc.2 · sdk-minimal · Flash3,1611,920 / 60.7%565$0.000531–$0.001062
DSH 0.2.0-rc.2 · headless · Flash53,47039,808 / 74.4%755$0.002622–$0.005243
Claude Code 2.1.226 · local config · Flash191,383105,088 / 54.9%851$0.013770–$0.027540
Claude Code 2.1.226 · --bare · Flash9,4116,656 / 70.7%709$0.000859–$0.001717
Codex CLI 0.154.0 · GPT-5.6 Sol109,45481,024 / 74.0%1,074Unknown / different billing
Hermes 0.19.0 · deepseek-chat alias22,23517,408 / 78.3%436Unknown / different billing

04 / A higher hit rate can still cost more

DSH headless had a 74.4% cache-read share; sdk-minimal had 60.7%. Yet the headless run's estimated peak cost was about $0.005243 versus $0.001062, because it sent far more total input and generated more output. A high hit rate can be the result of repeatedly re-sending a large stable prefix. That prefix is cheaper to read than to process cold, but it still has a price.

Our Save Tokens formula works once you put the actual counts into its terms. For providers with a separate cache-write tariff, add cache-write tokens × their rate. Also count billed tools, retries, and failed attempts. Price per token and cache percentage alone cannot answer cost per successful task.

DeepSeek Flash pricing and served-model alias, checked October 6, 2026 ↗Anthropic: separate cache reads, writes and uncached input ↗

Estimated model cost = (uncached input × miss price + cached input × hit price + cache writes × write price + output × output price) / 1,000,000

05 / The 'built-in prompt' is only part of the payload

An agent request can contain base instructions, tool names and schemas, descriptions of available skills, project files, prior messages and tool results. Pi documents that it assembles a request from its system prompt, active session, tools, discovered context files and skill descriptions. A single total-input number cannot tell us how many tokens came from any one component. Our agent logs provide usage, not a complete, comparable decomposition of every model-facing payload.

One independent payload inspection by cd4n1 is instructive: for a cold 'hello' task using the same Bedrock model, the author reported 28,407 input tokens for Claude Code, 12,374 for OpenCode and 2,768 for Pi. Their inspection attributed more of Claude Code's payload to tool definitions than to its system prompt. Their later cache calculation also showed why a cold first-turn ratio does not predict a warm session's bill. Those are the author's May 2026 versions, provider and setup, not numbers transferable to our October runs.

Another first-person Pi test reported 1,358 first-turn context tokens with optional resources disabled and 10,787 with 48 discovered skills. That same post used different models for its Pi versus Claude Code task comparison, so it cannot isolate an agent effect across products. Both reports support inspecting actual tools and loaded resources before calling a prompt 'small' or 'large'.

Pi context assembly ↗cd4n1's request-payload inspection ↗Bruce's Pi configuration test ↗

06 / Cache design is an agent design choice

A provider generally reuses the stable beginning of a request. Stable tool definitions and instructions can therefore earn repeated cache reads, while a changing tool list, timestamp, per-run marker or early context change can shorten the reusable prefix. Cache rules, minimum eligible lengths, retention and cache-write prices differ by provider and model; check the relevant API's usage fields rather than assuming a hit.

Cursor's engineering write-up describes moving infrequently used tool definitions out of static context and placing cache boundaries after stable layers. The team reports lower usage in its own A/B tests. This supports measuring tool selection and prefix placement, but its product-wide results do not predict a saving for a different agent or model. The optimization target is a correct result at acceptable total cost, with the tools and safeguards your task requires.

Cursor's own tool-loading and cache-boundary measurements ↗OpenAI cache semantics and usage fields ↗Anthropic cache semantics and tool definitions ↗

07 / How to test your own agent setup

Freeze a representative set of tasks and a written acceptance test before measuring. For each agent, record the exact version, model alias or snapshot, provider, profile, enabled tools, skills and project instructions. Run both a cold session and a realistic multi-step session. Save uncached input, cache reads, cache writes, output, tool charges, retries, test passes and wall-clock duration for every attempt. Compare a lean and a normal configuration of the same agent before blaming the model.

This pilot has only one easy task and one successful run per configuration. Tool choices and number of model calls varied. The regular Claude run reflects existing local customization, not a clean default installation; Pi was deliberately stripped down; DSH sdk-minimal is an SDK profile, not the Web UI's Minimal preset. Provider cache state was not reset. Codex used a different model and billing path, and Hermes's normalized alias could not be tied to a verified Flash price. None of these observations establishes a general cheapest-agent ranking or a prompt-size ranking.