Caching guide
When prompt caching costs more than it saves
Cache writes cost more than ordinary input on Claude and newer GPT models. Worked examples show when a write never pays back and caching adds cost.
· Facts checked October 3, 2026
- Claude Code
- Codex
- Gemini CLI
The short answer
Prompt caching costs more than it saves when a cache write is not read enough times before it expires. On Anthropic models a 5-minute write costs 1.25 times the input price and a 1-hour write costs 2 times, so a write that is never read is pure overhead, while a single read repays a 5-minute write and two reads repay a 1-hour write.
At a glance
| Anthropic writes | 1.25x input for 5 minutes, 2x input for 1 hour. Reads cost 0.1x on most models. |
|---|---|
| OpenAI writes | 1.25x input from GPT-5.6 on; earlier models add no write charge. |
| Implicit caching adds no write price; explicit caching adds an hourly storage charge. | |
| Break-even | One read for a 5-minute write, two reads for a 1-hour write (Anthropic's rule of thumb). |
| Common cause | Pauses longer than the cache lifetime, and switching models mid-session. |
A worked example at Claude Sonnet 5 rates
Take a 100,000-token context at Claude Sonnet 5 prices: $2 per million input tokens, $2.50 for a 5-minute cache write, $4 for a 1-hour write, and $0.20 for a cache read. Sent once as ordinary input, the context costs $0.20.
Written to a 5-minute cache, the same tokens cost $0.25. If nothing reads them before they expire, caching cost $0.05 more than not caching. One read adds $0.02 instead of another $0.20, so the cached total of $0.27 beats $0.40 uncached, and every further read widens the gap.
A 1-hour write costs $0.40. With no reads that is $0.20 of overhead, and with one read the total is $0.42 against $0.40 uncached, still a loss. Only the second read moves it ahead, at $0.44 against $0.60.
How a coding session ends up paying the premium
Agentic tools resend the whole conversation on every turn, so in a steady session almost every token is a cache read and caching saves a great deal. The overhead appears at the edges. Step away for longer than the cache lifetime and the next request writes the whole context again at the premium price.
Switching models has the same effect, because each model keeps its own cache. So does a context change near the start of the prompt, which invalidates everything after it. On a Claude subscription, Claude Code uses the 1-hour cache for the main conversation, which tolerates pauses but raises the price of a write that is never reused.
How to see it in your own usage
The free cache savings calculator shows the net result for any mix of writes, reads, and ordinary input. everyaitoken does the same from your real history: it prices cache reads, 5-minute writes, and 1-hour writes at each provider's rates and reports net savings, or cache overhead when the result is negative. For choosing between the two write lifetimes, see 5-minute vs 1-hour cache writes.
Your own numbers
See what your AI coding really costs.
everyaitoken reads Claude Code, Codex, Cursor, OpenRouter, Gemini CLI, and OpenCode on your Mac. It shows limits with reset times, API-equivalent cost, and what caching saved or cost, in the menu bar and 17 widgets. $9 once.
FAQ
Questions
Can prompt caching increase costs?
Yes. When cache writes cost more than ordinary input, as on Anthropic models and on OpenAI models from GPT-5.6 on, a write that is never read or read too few times costs more than sending the tokens uncached.
How many cache reads does a write need to pay off?
On Anthropic models, one read repays a 5-minute write at 1.25x input, and two reads repay a 1-hour write at 2x input. The cache hit rate guide turns this into a percentage.
Does Gemini charge for cache writes?
Google publishes no separate write price for implicit caching, so written tokens cost ordinary input. Explicit caching adds a storage charge for every hour the cache lives.
Why does my cache keep getting rewritten?
Usually because of pauses longer than the cache lifetime, a model switch, or a change early in the prompt. Each one forces the next request to write the context again.
Sources
Checked October 3, 2026. Plans and limits change; the linked pages are the authority.
