Skip to content

Caching guide

What is a good cache hit rate for AI coding?

How to read a prompt cache hit rate, the break-even rate for Anthropic's 5-minute and 1-hour writes, and what a falling hit rate usually means.

· Facts checked October 3, 2026

  • Claude Code
  • Codex
  • Gemini CLI

The short answer

A cache hit rate is the share of input tokens served from the cache. With Anthropic's prices, caching breaks even at roughly 22% for 5-minute writes and 53% for 1-hour writes, ignoring ordinary input, and an agentic coding session that resends its conversation every turn normally sits far above both.

At a glance

DefinitionCache read tokens divided by all input-side tokens: reads, writes, and ordinary input.
Break-even, 5-minuteAbout 22% of written plus read tokens on models with 0.1x reads.
Break-even, 1-hourAbout 53% on the same models.
Warning signA hit rate that drops while usage stays steady, which usually means expired or rebuilt caches.

Where the break-even numbers come from

On Anthropic models with 0.1x reads, every read saves 0.9 times the input price, and every 5-minute write costs an extra 0.25 times. Caching pays once reads times 0.9 exceed writes times 0.25, which works out to about 0.28 reads per written token, or a hit rate near 22% of the cached tokens.

A 1-hour write costs an extra full input price, so reads must exceed about 1.1 per written token, a hit rate near 53%. On models with cheaper reads, such as Claude Opus 5.5 at 0.05x input, the thresholds are slightly lower. These figures leave out ordinary uncached input, which costs the same with or without caching.

What a healthy coding session looks like

Coding agents send the system prompt, tool definitions, files, and the entire conversation on every turn, then add a little new text at the end. After the first write, nearly all of each request is a read, so steady sessions clear the break-even line by a wide margin.

The number is more useful as a trend than as a target. A hit rate that falls while your work looks the same usually means caches are expiring between requests, models are being switched, or something early in the prompt changes each turn. The 5-minute vs 1-hour guide explains the expiry side.

Reading it alongside the dollar figure

A high hit rate with negative savings is possible when expensive 1-hour writes are spread over too few reads, so read the hit rate together with net savings. The cache savings calculator computes both from any token mix, and everyaitoken shows both per day, tool, and model from your own history.

Your own numbers

See what your AI coding really costs.

everyaitoken reads Claude Code, Codex, Cursor, OpenRouter, Gemini CLI, and OpenCode on your Mac. It shows limits with reset times, API-equivalent cost, and what caching saved or cost, in the menu bar and 17 widgets. $9 once.

FAQ

Questions

How is cache hit rate calculated?

Divide cache read tokens by all input-side tokens, meaning reads, cache writes, and ordinary input together. Output tokens are not part of it.

What cache hit rate do I need to save money?

With Anthropic's 0.1x reads, about 22% of cached tokens for 5-minute writes and about 53% for 1-hour writes. Below that, the write premium outweighs the read discount, as the caching overhead guide shows.

Why did my cache hit rate drop?

Common causes are long pauses that let caches expire, switching models, and edits near the start of the prompt, each of which forces a fresh write.

Is a 90% cache hit rate good?

Yes. It is far above the break-even point for both 5-minute and 1-hour writes. At that level, how much caching saves depends more on how often the cache is rewritten after pauses than on the hit rate itself.

Sources

Checked October 3, 2026. Plans and limits change; the linked pages are the authority.

  1. Anthropic: Prompt caching
  2. OpenAI: Prompt caching
  3. Google: Context caching