01 / Caching reduces input processing
Prompt caching reuses a previously processed input prefix. It does not mean the answer is cached or that output tokens are free. Separate cached input, uncached input, and output when comparing costs.
DeepSeek context caching: matching prefixes and output behavior ↗
02 / Keep the stable prefix first
Put stable instructions, tool definitions, and repeated reference material first. Put the current question and changing information last. Timestamps, random IDs, reordered tools, and rewritten instructions can reduce prefix reuse.
[Stable] Instructions → Tools → Reference material
[Changing] Current question → New context03 / Measure cache hits from API usage
Check the provider's usage fields instead of assuming repeated requests hit the cache. The first request may establish a cached prefix; later requests are not guaranteed to reuse all of it. This is an illustrative response, not a measured result.
DeepSeek API: cache hit and miss token fields ↗
{"usage":{"prompt_cache_hit_tokens":16000,"prompt_cache_miss_tokens":2234,"completion_tokens":412}}04 / Verify with your own workload
Keep the model and prefix fixed, then vary only the final question. Record cache hits, total cost, and answer quality. Compare with your previous input format. Retention, minimum lengths, and write charges vary by provider.
Sources and further reading
DS-V4-HUUB · DeepSeek Context Caching & Cache-Hit Rules ↗api-docs.deepseek.com · Official documentation ↗Adapted from the source guide, last updated 2026-06-26. Savings depend on your workload and must be measured.
