Model comparison
Kimi K3 or GPT-6 Sol: which costs less for agentic coding?
GPT-6 Sol undercuts Kimi K3 on every rate, and an agentic coding session costs $2.10 against $2.85. Both list 1.05M context, with different limits inside.
· Prices as of September 28, 2026
Kimi K3
Moonshot AI · Released July 16, 2026
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Kimi K3 facts and comparisonsGPT-6 Sol
OpenAI · Released September 22, 2026
The mid-priced GPT-6 model, which OpenAI pitches for complex coding and agent workflows and which the Codex docs recommend for complex coding.
GPT-6 Sol facts and comparisons
The short answer
GPT-6 Sol is cheaper on every rate, and the example agentic coding session costs $2.10 on it against $2.85 on Kimi K3, a 26% saving. Pick GPT-6 Sol if you work in Codex, whose docs recommend it for complex coding; pick Kimi K3 for open weights or for responses longer than 128K tokens.
Choose Kimi K3 if
- You need long single responses: Kimi K3's output can be raised to its whole 1.05M window, against 128K on GPT-6 Sol.
- You want open weights, published under Moonshot's own Kimi K3 license.
- You send very long prompts, since GPT-6 Sol bills requests over 272K input tokens at 2x for input and cache, and the Kimi K3 rates here list no such tier.
Choose GPT-6 Sol if
- You want lower rates across the board: $2 input, $0.20 cache hits, and $10 output per million tokens, against $3, $0.30, and $15.
- Codex is your agent: it is built around OpenAI's models, and its docs name GPT-6 Sol for complex coding.
- You want explicit control over caching, with up to four breakpoints and a cached prefix that stays reusable for at least 30 minutes after its last use.
Side by side
Specs and prices
| Fact | Kimi K3 | GPT-6 Sol |
|---|---|---|
| Maker | Moonshot AI | OpenAI |
| API model id | kimi-k3 | gpt-6-sol |
| Released | July 16, 2026 | September 22, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 1.05M tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $3 | $2 |
| Cache hit, per 1M | $0.30 | $0.20 |
| Cache write, per 1M | $3 (same as input) | $2.50 |
| Output, per 1M tokens | $15 | $10 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. GPT-6 Sol: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Kimi K3 | GPT-6 Sol |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.85 | $2.10 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $1.29 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $313.50 | $231.00 |
| Where the session’s cost goes | ||
| Cache writes | $1.20 | $1.00 |
| Cache reads | $0.60 | $0.40 |
| Uncached input | $0.30 | $0.20 |
| Output | $0.75 | $0.50 |
- caching saves on the session with Kimi K3 (65%)
- $5.40
- caching saves on the session with GPT-6 Sol (62%)
- $3.40
GPT-6 Sol costs less on every line
GPT-6 Sol charges $2 per million input tokens, $0.20 per million cache hits, and $10 per million output tokens. Kimi K3 charges $3, $0.30, and $15, each 50% more. On uncached work the ratio carries straight through: the large one-off review costs $0.40 against $0.60, and the output-heavy generation $0.86 against $1.29.
Cache writes are where the gap narrows. OpenAI charges 1.25x input to write the cache, $2.50 per million on GPT-6 Sol. Kimi bills its 5-minute write at its input price, $3. OpenAI's premium still leaves its write price below Kimi's, but only 17% below rather than 33%.
In the example session that means $1.00 of writes on GPT-6 Sol against $1.20 on Kimi K3, plus $0.20 more for reads, $0.25 more for output, and $0.10 more for fresh input on Kimi K3. The session costs $2.10 against $2.85, or $231.00 against $313.50 over 110 sessions a month, a gap of $82.50.
The same 1.05M window with different limits inside it
Kimi K3 and GPT-6 Sol both list a 1.05M context window, but they use it differently. GPT-6 Sol accepts up to 922K input tokens and writes up to 128K of output. Kimi K3 defaults to 131,072 output tokens per request and can raise output all the way to its full window.
Pricing inside the window differs as well. Kimi K3's rates in this comparison list no long-context tier, whereas a GPT-6 Sol request past 272K input tokens is billed in full at 2x for input and cache and 1.5x for output. The example session stays under 200K per request, so neither rule changes its numbers, but an agent that packs a large repository into one prompt will cross GPT-6 Sol's line.
Reasoning settings change output length too. Medium is GPT-6 Sol's default effort, both in the API and in Codex. Output is 26% of the Kimi K3 session and 24% of the GPT-6 Sol session here, so a model that thinks longer on your tasks shifts the comparison.
How the Kimi and OpenAI caches behave
For GPT-6 Sol, OpenAI caches automatically and also takes up to four explicit breakpoints. Its cached prefixes stay reusable for at least 30 minutes after their last use, and caching begins at 1,024 input tokens. A hit costs 0.1x input.
Kimi's cache needs no setup: prefixes live 5 minutes by default or an hour on request, and every hit restarts that lifetime free of charge. A hit costs one-tenth of input, the same ratio as OpenAI. Kimi bills the 1-hour write at $6 per million, which this comparison does not use: the session prices every Kimi write at the 5-minute default.
Moonshot says coding workloads on its official API hit the Kimi K3 cache more than 90% of the time. In the example session, caching saves 65% of the uncached cost on Kimi K3 and 62% on GPT-6 Sol. The cache walkthrough on our homepage shows how the session's Kimi K3 and GPT-6 Sol writes and reads add up.
Where they run and what their makers claim
Codex, OpenAI's own agent, is built around OpenAI's models and points to GPT-6 Sol for complex coding in its docs. OpenCode, OpenRouter, and GitHub Copilot carry both models.
OpenAI describes GPT-6 Sol as "built to power complex coding and agentic workflows," and it is the mid-priced model of the GPT-6 family. Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," and aims it at long engineering tasks with minimal supervision, large codebases, and terminal tools.
Kimi K3 has open weights under Moonshot's own license, and GPT-6 Sol has none. The Kimi prices here are Moonshot's own API rates, and OpenRouter providers can charge differently. EveryToken prices GPT-6 Sol at OpenAI's rates from Codex and OpenCode history, and prices Kimi K3 from OpenRouter's catalog when you use it through OpenRouter.
Prompt caching
How each maker bills cached tokens
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what Kimi K3 and GPT-6 Sol really cost you.
everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GPT-6 Sol cheaper than Kimi K3?
Yes, on every rate. The example agentic session costs $2.10 on GPT-6 Sol against $2.85 on Kimi K3, and uncached work shows the full 1.5x list-price gap.
Does OpenAI charge for cache writes on GPT-6 Sol?
Yes, 1.25x the input price, which is $2.50 per million. Kimi K3's 5-minute write costs $3, the same as its input, so GPT-6 Sol still writes to the cache for less.
Which model has the larger context window?
Both list 1.05M tokens. GPT-6 Sol accepts up to 922K of that as input and writes up to 128K, while Kimi K3 can raise its output to the full window.
Is Kimi K3 in Codex's model list?
The sources behind this comparison don't say. Codex is OpenAI's own agent, built around and defaulting to OpenAI's models, and custom configuration can point it at other providers' compatible endpoints, which this page doesn't cover. OpenCode, OpenRouter, and GitHub Copilot offer both Kimi K3 and GPT-6 Sol.
Sources
- Kimi API: Chat pricing
- Kimi API: Kimi K3 quickstart
- Kimi: Kimi K3
- OpenRouter: Kimi K3
- OpenCode docs: Zen
- Cursor docs: Models
- GitHub Docs: Supported AI models in Copilot
- OpenAI: API pricing
- OpenAI docs: GPT-6 Sol
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Sol
- OpenCode docs: Zen
- OpenAI: Prompt caching