Estimate the cost before changing your workflow.

Estimated cost per request
=
Input (M tokens) × cache-hit rate × cached-input price
+Input (M tokens) × (1 − cache-hit rate) × uncached-input price
+Output (M tokens) × output price

Token usage is in millions; prices are per 1M tokens. Cache-hit rate is a fraction from 0 to 1.

Use your provider's actual rates and measured cache hits. Cache writes, tools, and other charges may add cost. Then check whether the change still produces an acceptable answer.

Estimate your token costs →

Token efficiency & context engineering

Original research and practical experience, with short notes to help you choose.