Skip to content

Model comparison

Kimi K3 vs DeepSeek-V4-Pro: why DeepSeek costs a third

An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.

· Prices as of September 28, 2026

  • Kimi K3

    Moonshot AI · Released July 16, 2026

    Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.

    Kimi K3 facts and comparisons
  • DeepSeek-V4-Pro

    DeepSeek · Released August 13, 2026

    DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.

    DeepSeek-V4-Pro facts and comparisons

The short answer

DeepSeek-V4-Pro costs a third as much as Kimi K3 on the example agentic coding session, $0.95 against $2.85, and its off-peak hours cut its rates by 50% again. Pick Kimi K3 if you want it inside Cursor or GitHub Copilot or need outputs beyond 384K tokens; pick DeepSeek-V4-Pro when cost leads, especially for cache-heavy agent loops.

Choose Kimi K3 if

  • You work in Cursor or GitHub Copilot, which offer Kimi K3 and do not offer DeepSeek-V4-Pro.
  • You need single responses beyond DeepSeek-V4-Pro's 384K limit, up to Kimi K3's full 1.05M window.
  • You want the option of a 1-hour cache lifetime, which Kimi offers at $6 per million tokens written.

Choose DeepSeek-V4-Pro if

  • Cost leads: DeepSeek-V4-Pro lists $1.32 input and $3.96 output per million tokens at peak, against $3 and $15.
  • Your agent loops reread a large context, where a DeepSeek-V4-Pro hit at $0.044 per million costs 85% less than Kimi K3's $0.30.
  • You can run jobs off-peak, when DeepSeek charges 50% less.
  • You want open weights under the MIT license.

Side by side

Specs and prices

FactKimi K3DeepSeek-V4-Pro
MakerMoonshot AIDeepSeek
API model idkimi-k3deepseek-v4-pro
ReleasedJuly 16, 2026August 13, 2026
StatusCurrentCurrent
Context window1.05M tokens1M tokens
Max output1.05M tokens384K tokens
Open weightsYesYes
Input, per 1M tokens$3$1.32
Cache hit, per 1M$0.30$0.044
Cache write, per 1M$3 (same as input)$1.32 (same as input)
Output, per 1M tokens$15$3.96
Runs inCursor, OpenCode, OpenRouter, and GitHub CopilotOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadKimi K3DeepSeek-V4-Pro
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.85$0.95
Large one-off review, 150K input with no cache hits, 10K output$0.60$0.24
Output-heavy generation, 30K input, 80K output$1.29$0.36
A month of sessions, 110 sessions: 5 a day, 22 working days$313.50$104.06
Where the session’s cost goes
Cache writes$1.20$0.53
Cache reads$0.60$0.09
Uncached input$0.30$0.13
Output$0.75$0.20
caching saves on the session with Kimi K3 (65%)
$5.40
caching saves on the session with DeepSeek-V4-Pro (73%)
$2.55

Where DeepSeek-V4-Pro's 3x advantage comes from

Kimi K3 lists $3 input and $15 output per million tokens. DeepSeek-V4-Pro lists $1.32 and $3.96 at its peak rates, so input is 2.3x dearer on Kimi K3 and output 3.8x. The large one-off review costs $0.60 against $0.24, and the output-heavy generation $1.29 against $0.36.

The agentic session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro, $1.90 apart. Every line contributes: $0.67 from cache writes, $0.55 from output, $0.51 from cache reads, and $0.17 from fresh input. Over 110 sessions a month, that comes to $313.50 against $104.06.

Those DeepSeek figures are peak prices. DeepSeek charges 50% less off-peak, and peak covers only 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Work that runs outside those hours widens the gap further, since the Kimi K3 figures are its standard list prices.

Two approaches to prompt caching

DeepSeek's disk cache is on by default for every account, with no code changes and no fee for writing it. A hit on V4-Pro costs $0.044 per million, 3.3% of input, or 96.7% off.

Moonshot's cache for Kimi K3 is also automatic but works in lifetimes: 5 minutes by default, 1 hour when asked, with each hit extending it at no charge. Writes are billed on their own, at $3 per million for 5 minutes, equal to input, and $6 for an hour. A hit costs $0.30, one-tenth of input.

The result is a 6.8x gap on the rate a long agent loop uses most. In the session, 2M cached tokens cost $0.60 on Kimi K3 and $0.09 on DeepSeek-V4-Pro. Caching saves 73% on DeepSeek-V4-Pro and 65% on Kimi K3 against the uncached cost, and Moonshot says Kimi K3 coding traffic on its API hits the cache over 90% of the time. The homepage cache example explains the session.

Output limits, tools, and open weights

Both take about 1M tokens of context, 1.05M on Kimi K3 and 1M on DeepSeek-V4-Pro. Both output limits are large: DeepSeek-V4-Pro writes up to 384K per request, and Kimi K3 defaults to 131,072 but can raise output to its full 1.05M window.

Kimi K3 is offered in Cursor, OpenCode, OpenRouter, and GitHub Copilot. DeepSeek-V4-Pro is offered through OpenRouter and in OpenCode, and DeepSeek says it added native support for OpenAI's Responses API, adapted for Codex with one-click setup.

Both ship open weights, DeepSeek under the MIT license and Moonshot under its own Kimi K3 license, and this page does not estimate self-hosting costs. The Kimi K3 and DeepSeek-V4-Pro prices here are each maker's own, and OpenRouter routes open-weight models to providers whose prices can differ. EveryToken prices Kimi K3 and DeepSeek-V4-Pro from OpenRouter's catalog when either runs through OpenRouter.

What Moonshot and DeepSeek say, and what comes next

Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," aimed at long engineering tasks with minimal supervision, large codebases, and terminal tools, and at frontend and game work that mixes code with visual reasoning from screenshots.

DeepSeek says its general-availability release "greatly enhances agent capabilities, with particularly significant performance improvements in production environments." It offers reasoning effort at low, high, and max.

DeepSeek-V4-Pro is also a model in transition. DeepSeek says service continues with billing unchanged until a V4.1 Pro model arrives, and says its smaller DeepSeek-V4.1-Flash already outperforms V4-Pro. Kimi K3, released on July 16, 2026, is Moonshot's current flagship.

Prompt caching

How each maker bills cached tokens

Moonshot AI

The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.

Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.

Source: Kimi API: Chat pricing

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

Your own numbers

See what Kimi K3 and DeepSeek-V4-Pro really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is DeepSeek-V4-Pro than Kimi K3?

About 3x on the example agentic session, $0.95 against $2.85, at DeepSeek's peak rates. Off-peak, DeepSeek charges 50% less. The gap is widest on cache hits, $0.044 against $0.30 per million.

Can I use DeepSeek-V4-Pro in Cursor or GitHub Copilot?

No. Neither offers DeepSeek-V4-Pro, while both offer Kimi K3. DeepSeek-V4-Pro is available through OpenRouter and in OpenCode.

Does DeepSeek charge for cache writes?

No. DeepSeek lists no fee for writing its disk cache, so written tokens cost the $1.32 input rate. Kimi bills its 5-minute writes at $3 per million, the same as its input, and 1-hour writes at $6.

Is a newer DeepSeek Pro model coming?

Until a V4.1 Pro model arrives, DeepSeek says, V4-Pro service continues with billing unchanged. The current snapshot is DeepSeek-V4-Pro-0813.

  • DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro

    DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

  • DeepSeek-V4-Pro vs Claude Opus 5.5

    Claude Opus 5.5 costs $4.40 for a cached coding session that costs $0.95 on DeepSeek-V4-Pro. Cache writes and output drive the gap; a V4.1 Pro is planned.

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • DeepSeek-V4-Pro vs GPT-6 Sol

    DeepSeek-V4-Pro costs $0.95 for a cached coding session against $2.10 on GPT-6 Sol. How the gap forms, and what DeepSeek's Codex support does and doesn't mean.

  • Grok 4.7 vs Kimi K3

    Grok 4.7 and Kimi K3 both run in Cursor and GitHub Copilot. Grok 4.7 costs less on every workload, while Kimi K3 adds open weights and a 1.05M window.

  • DeepSeek-V4-Pro vs Claude Sonnet 5

    DeepSeek-V4-Pro costs $0.95 for a cached coding session that costs $2.40 on Claude Sonnet 5, and cache writes explain most of it. Rates, limits, and tools.