Skip to content

Glossary

What is a cache write premium?

A cache write premium is the extra a provider charges to store a prompt in its cache, above the ordinary input price. Rates for Anthropic, OpenAI, and Google.

· Facts checked October 3, 2026

The short answer

A cache write premium is the amount above the ordinary input price that a provider charges to write a prompt into its cache. Anthropic charges 1.25 times input for a 5-minute write and 2 times for a 1-hour write, and OpenAI charges 1.25 times from GPT-5.6 on, so the premium is 25% or 100% of the input price.

At a glance

Anthropic5-minute write 1.25x input, 1-hour write 2x input.
OpenAI1.25x input from GPT-5.6 on; no premium on GPT-5.5 and earlier.
GoogleNo write price for implicit caching; explicit caching adds hourly storage.
Paid back byCache reads, which cost 0.1x input on most of these models.

Why providers charge it

Writing a prompt to the cache stores the model's processed state so later requests can reuse it. The provider prices that storage into the first request, then discounts every request that reads it. The premium is the cost of the bet that the prompt will be reused.

With Claude prices, a 5-minute write on Claude Opus 5.5 costs $5 per million tokens against $4 for ordinary input, a $1 premium, while each read costs $0.20 instead of $4.

When the premium becomes overhead

If the cached prompt expires before it is read enough times, the premium is never recovered. That is the case covered in when prompt caching costs more. Tools that report cache savings without subtracting write premiums overstate what caching saved, which is why everyaitoken reports savings net of premiums and shows overhead when the net is negative.

Your own numbers

See what your AI coding really costs.

everyaitoken reads Claude Code, Codex, Cursor, OpenRouter, Gemini CLI, and OpenCode on your Mac. It shows limits with reset times, API-equivalent cost, and what caching saved or cost, in the menu bar and 17 widgets. $9 once.

FAQ

Questions

How much is Anthropic's cache write premium?

25% above the input price for a 5-minute write, which costs 1.25 times input, and 100% above it for a 1-hour write, which costs 2 times input.

Does OpenAI charge for cache writes?

From GPT-5.6 on, a cache write costs 1.25 times the uncached input price. GPT-5.5 and earlier bill written tokens as ordinary input with no premium.

Is the write premium worth paying?

Usually, in coding sessions that reuse the same context many times. One read repays a 5-minute write on Anthropic models; a 1-hour write needs two, as the 5-minute vs 1-hour guide shows.

Sources

Checked October 3, 2026. Plans and limits change; the linked pages are the authority.

  1. Anthropic: Prompt caching
  2. OpenAI: Prompt caching
  3. Google: Context caching