Free developer tool
What is your prompt cache really saving?
Compare the same token usage with and without caching. Include the write premium, see the net result, and share a reproducible scenario. No account required.
Your tokens. Your rates.
Start with this illustrative example, then enter your model’s rates and token volumes for the same period. Rates are editable estimates in USD per million tokens, not live provider prices.
Sharing adds these values to the link. Use illustrative values if your usage or pricing is confidential.
Estimated net savings
$4.65
56.4% less expensive than the no-cache estimate.
- Without caching
- $8.25
- With caching
- $3.60
- Savings from cache reads
- $5.40
- Cache write premium
- $0.75
- Cache hit rate
- 80%
Net savings subtract the write premium from the savings on reads. If a cache is written and seldom reused, caching can increase your cost.
Hit rate is cache-read tokens divided by all input-side tokens, including ordinary input and both kinds of writes. Write premium compares each write rate with the ordinary input rate; a negative premium means writes are cheaper.
This is an API token-cost estimate, not your subscription bill. It excludes cache storage fees, taxes, batch discounts, negotiated pricing, and other charges. Token categories should not overlap. Calculations retain full precision; only displayed values are rounded.
How the calculation works
Without caching
All ordinary input, cache write, and cache read tokens are priced as ordinary input. Output uses the output rate.
With caching
Each token category uses its own rate. Net savings equal the uncached estimate minus the cached estimate. A negative result means a cost increase.
Cache hit rate is read tokens divided by all input-side tokens. The calculator keeps full numeric precision until display. The example rates are illustrative and editable; check your provider’s current model pricing before using the result.
Provider billing rules differ. Some charge separate cache storage fees; those are outside this token-only estimate. Enter actual cache write rates, not just the extra premium, and leave unused categories at zero.
Read Anthropic’s cache write and read methodologyQuestions about the estimate
- Can prompt caching cost more than ordinary input?
- Yes. If cache writes cost more than ordinary input and too few tokens are read from the cache, the write premium can exceed the read savings. This calculator shows a cost increase when that happens.
- Does this calculate my Claude Code or ChatGPT subscription bill?
- No. This is an API-equivalent token cost estimate. It does not predict subscription limits, subscription charges, taxes, storage charges, batch discounts, or regional pricing adjustments.
- What should I enter as ordinary input?
- Enter only input tokens that were neither cache writes nor cache reads. Each input token belongs in exactly one category. Output tokens are counted separately and priced identically in both scenarios.
See the numbers from your own coding tools
everyaitoken brings supported AI coding usage and cache cost accounting into a native Mac app, with multiple accounts and models in one place. The calculator above is free and works independently of the app.
Explore everyaitoken