Model comparison
Kimi K3 vs DeepSeek-V4-Pro: why DeepSeek costs a third
An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.
· Prices as of September 28, 2026
Kimi K3
Moonshot AI · Released July 16, 2026
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Kimi K3 facts and comparisonsDeepSeek-V4-Pro
DeepSeek · Released August 13, 2026
DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.
DeepSeek-V4-Pro facts and comparisons
The short answer
DeepSeek-V4-Pro costs a third as much as Kimi K3 on the example agentic coding session, $0.95 against $2.85, and its off-peak hours cut its rates by 50% again. Pick Kimi K3 if you want it inside Cursor or GitHub Copilot or need outputs beyond 384K tokens; pick DeepSeek-V4-Pro when cost leads, especially for cache-heavy agent loops.
Choose Kimi K3 if
- You work in Cursor or GitHub Copilot, which offer Kimi K3 and do not offer DeepSeek-V4-Pro.
- You need single responses beyond DeepSeek-V4-Pro's 384K limit, up to Kimi K3's full 1.05M window.
- You want the option of a 1-hour cache lifetime, which Kimi offers at $6 per million tokens written.
Choose DeepSeek-V4-Pro if
- Cost leads: DeepSeek-V4-Pro lists $1.32 input and $3.96 output per million tokens at peak, against $3 and $15.
- Your agent loops reread a large context, where a DeepSeek-V4-Pro hit at $0.044 per million costs 85% less than Kimi K3's $0.30.
- You can run jobs off-peak, when DeepSeek charges 50% less.
- You want open weights under the MIT license.
Side by side
Specs and prices
| Fact | Kimi K3 | DeepSeek-V4-Pro |
|---|---|---|
| Maker | Moonshot AI | DeepSeek |
| API model id | kimi-k3 | deepseek-v4-pro |
| Released | July 16, 2026 | August 13, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1M tokens |
| Max output | 1.05M tokens | 384K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $3 | $1.32 |
| Cache hit, per 1M | $0.30 | $0.044 |
| Cache write, per 1M | $3 (same as input) | $1.32 (same as input) |
| Output, per 1M tokens | $15 | $3.96 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Kimi K3 | DeepSeek-V4-Pro |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.85 | $0.95 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 | $0.24 |
| Output-heavy generation, 30K input, 80K output | $1.29 | $0.36 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $313.50 | $104.06 |
| Where the session’s cost goes | ||
| Cache writes | $1.20 | $0.53 |
| Cache reads | $0.60 | $0.09 |
| Uncached input | $0.30 | $0.13 |
| Output | $0.75 | $0.20 |
- caching saves on the session with Kimi K3 (65%)
- $5.40
- caching saves on the session with DeepSeek-V4-Pro (73%)
- $2.55
Where DeepSeek-V4-Pro's 3x advantage comes from
Kimi K3 lists $3 input and $15 output per million tokens. DeepSeek-V4-Pro lists $1.32 and $3.96 at its peak rates, so input is 2.3x dearer on Kimi K3 and output 3.8x. The large one-off review costs $0.60 against $0.24, and the output-heavy generation $1.29 against $0.36.
The agentic session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro, $1.90 apart. Every line contributes: $0.67 from cache writes, $0.55 from output, $0.51 from cache reads, and $0.17 from fresh input. Over 110 sessions a month, that comes to $313.50 against $104.06.
Those DeepSeek figures are peak prices. DeepSeek charges 50% less off-peak, and peak covers only 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Work that runs outside those hours widens the gap further, since the Kimi K3 figures are its standard list prices.
Two approaches to prompt caching
DeepSeek's disk cache is on by default for every account, with no code changes and no fee for writing it. A hit on V4-Pro costs $0.044 per million, 3.3% of input, or 96.7% off.
Moonshot's cache for Kimi K3 is also automatic but works in lifetimes: 5 minutes by default, 1 hour when asked, with each hit extending it at no charge. Writes are billed on their own, at $3 per million for 5 minutes, equal to input, and $6 for an hour. A hit costs $0.30, one-tenth of input.
The result is a 6.8x gap on the rate a long agent loop uses most. In the session, 2M cached tokens cost $0.60 on Kimi K3 and $0.09 on DeepSeek-V4-Pro. Caching saves 73% on DeepSeek-V4-Pro and 65% on Kimi K3 against the uncached cost, and Moonshot says Kimi K3 coding traffic on its API hits the cache over 90% of the time. The homepage cache example explains the session.
Output limits, tools, and open weights
Both take about 1M tokens of context, 1.05M on Kimi K3 and 1M on DeepSeek-V4-Pro. Both output limits are large: DeepSeek-V4-Pro writes up to 384K per request, and Kimi K3 defaults to 131,072 but can raise output to its full 1.05M window.
Kimi K3 is offered in Cursor, OpenCode, OpenRouter, and GitHub Copilot. DeepSeek-V4-Pro is offered through OpenRouter and in OpenCode, and DeepSeek says it added native support for OpenAI's Responses API, adapted for Codex with one-click setup.
Both ship open weights, DeepSeek under the MIT license and Moonshot under its own Kimi K3 license, and this page does not estimate self-hosting costs. The Kimi K3 and DeepSeek-V4-Pro prices here are each maker's own, and OpenRouter routes open-weight models to providers whose prices can differ. EveryToken prices Kimi K3 and DeepSeek-V4-Pro from OpenRouter's catalog when either runs through OpenRouter.
What Moonshot and DeepSeek say, and what comes next
Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," aimed at long engineering tasks with minimal supervision, large codebases, and terminal tools, and at frontend and game work that mixes code with visual reasoning from screenshots.
DeepSeek says its general-availability release "greatly enhances agent capabilities, with particularly significant performance improvements in production environments." It offers reasoning effort at low, high, and max.
DeepSeek-V4-Pro is also a model in transition. DeepSeek says service continues with billing unchanged until a V4.1 Pro model arrives, and says its smaller DeepSeek-V4.1-Flash already outperforms V4-Pro. Kimi K3, released on July 16, 2026, is Moonshot's current flagship.
Prompt caching
How each maker bills cached tokens
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
Your own numbers
See what Kimi K3 and DeepSeek-V4-Pro really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is DeepSeek-V4-Pro than Kimi K3?
About 3x on the example agentic session, $0.95 against $2.85, at DeepSeek's peak rates. Off-peak, DeepSeek charges 50% less. The gap is widest on cache hits, $0.044 against $0.30 per million.
Can I use DeepSeek-V4-Pro in Cursor or GitHub Copilot?
No. Neither offers DeepSeek-V4-Pro, while both offer Kimi K3. DeepSeek-V4-Pro is available through OpenRouter and in OpenCode.
Does DeepSeek charge for cache writes?
No. DeepSeek lists no fee for writing its disk cache, so written tokens cost the $1.32 input rate. Kimi bills its 5-minute writes at $3 per million, the same as its input, and 1-hour writes at $6.
Is a newer DeepSeek Pro model coming?
Until a V4.1 Pro model arrives, DeepSeek says, V4-Pro service continues with billing unchanged. The current snapshot is DeepSeek-V4-Pro-0813.
Sources
- Kimi API: Chat pricing
- Kimi API: Kimi K3 quickstart
- Kimi: Kimi K3
- OpenRouter: Kimi K3
- OpenCode docs: Zen
- Cursor docs: Models
- GitHub Docs: Supported AI models in Copilot
- DeepSeek API: Models and pricing
- DeepSeek API: Change log
- DeepSeek: V4 Pro release
- OpenRouter: DeepSeek-V4-Pro-0813
- DeepSeek API: Context caching