Model comparison
Kimi K3 vs GPT-6 Astra: two flagships, a 3.7x cost gap
GPT-6 Astra costs 3.3x Kimi K3 per token and 3.7x on an agentic coding session, $10.50 against $2.85. What each maker claims for its top model.
· Prices as of September 28, 2026
Kimi K3
Moonshot AI · Released July 16, 2026
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Kimi K3 facts and comparisonsGPT-6 Astra
OpenAI · Released September 3, 2026
OpenAI's top GPT-6 model and its most expensive, aimed at the hardest long-running work that spans many tools, including coding.
GPT-6 Astra facts and comparisons
The short answer
Kimi K3 is far cheaper: the example agentic coding session costs $2.85 on it against $10.50 on GPT-6 Astra, and Astra's input, output, and cache-hit rates are all 3.3x higher. Pick GPT-6 Astra if you work in Codex, where it is the default in the CLI's bundled model list, and you accept OpenAI's claim that it finishes tasks with fewer output tokens; pick Kimi K3 for cost, open weights, and outputs beyond 128K.
Choose Kimi K3 if
- Cost matters: at 110 sessions a month the example session totals $313.50 on Kimi K3 against $1,155.00 on GPT-6 Astra.
- You need responses longer than 128K tokens, since Kimi K3 can raise its output to its full 1.05M window.
- You want open weights under Moonshot's Kimi K3 license.
Choose GPT-6 Astra if
- You work in Codex, where GPT-6 Astra is the default in the CLI's bundled model list and starts at low reasoning effort.
- You judge models by cost per task rather than per token, and OpenAI's claim that Astra reaches stronger results with substantially fewer output tokens fits your work.
- You run long jobs across many tools, where OpenAI says Astra stays coherent better than GPT-5.6 Sol and earlier models.
Side by side
Specs and prices
| Fact | Kimi K3 | GPT-6 Astra |
|---|---|---|
| Maker | Moonshot AI | OpenAI |
| API model id | kimi-k3 | gpt-6-astra |
| Released | July 16, 2026 | September 3, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 1.05M tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $3 | $10 |
| Cache hit, per 1M | $0.30 | $1 |
| Cache write, per 1M | $3 (same as input) | $12.50 |
| Output, per 1M tokens | $15 | $50 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. GPT-6 Astra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Kimi K3 | GPT-6 Astra |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.85 | $10.50 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 | $2.00 |
| Output-heavy generation, 30K input, 80K output | $1.29 | $4.30 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $313.50 | $1,155.00 |
| Where the session’s cost goes | ||
| Cache writes | $1.20 | $5.00 |
| Cache reads | $0.60 | $2.00 |
| Uncached input | $0.30 | $1.00 |
| Output | $0.75 | $2.50 |
- caching saves on the session with Kimi K3 (65%)
- $5.40
- caching saves on the session with GPT-6 Astra (62%)
- $17.00
How big is the price gap between Kimi K3 and GPT-6 Astra?
GPT-6 Astra is OpenAI's most expensive GPT-6 model: $10 per million input tokens, $1 per million cache hits, and $50 per million output tokens. Kimi K3 charges $3, $0.30, and $15. Each of those rates is 3.3x higher on Astra, so the large one-off review costs $2.00 against $0.60 and the output-heavy generation $4.30 against $1.29.
The agentic coding session widens the gap to 3.7x, $10.50 against $2.85. Cache writes are the reason. OpenAI charges 1.25x input to write the cache, $12.50 per million on Astra, while Kimi's 5-minute write costs its $3 input rate. On the session's 400K written tokens that is $5.00 against $1.20, a $3.80 difference and half of the $7.65 total.
At 110 sessions a month, the session comes to $1,155.00 on GPT-6 Astra and $313.50 on Kimi K3, a difference of $841.50. These are API-equivalent estimates at each maker's own rates. ChatGPT plans that include Codex are priced differently, and OpenRouter providers serving Kimi K3's open weights can charge differently from Moonshot.
Cost per token versus cost per task
OpenAI's case for Astra rests partly on efficiency. It says Astra reaches stronger results with substantially fewer output tokens in several evaluations, for a lower estimated cost per task. If that holds on your work, the real gap is smaller than the per-token gap, because output carries the highest rate on both models.
The example workloads here hold token counts fixed on both models, so they cannot show that effect. Output is 24% of Astra's session cost and 26% of Kimi K3's. The cache writes and reads, which make up most of each session, are priced 3.3x to 4.2x higher on Astra whatever the output length.
Reasoning effort matters too. Codex starts Astra at low effort, which tends to keep output down. Moonshot's and OpenAI's tokenizers differ as well, so the same repository will not count as the same number of tokens on both. Compare Astra and Kimi K3 on a few real tasks and look at the cost of each finished change.
Caching and context on both models
Both models bill a cache hit at one-tenth of input: $1 per million on Astra and $0.30 on Kimi K3. For Astra, OpenAI adds up to four explicit breakpoints and keeps a cached prefix reusable for at least 30 minutes after its last use. Kimi caches for 5 minutes by default or 1 hour on request and bills a 1-hour write at $6 per million, while this comparison prices Kimi's writes at the 5-minute default.
Caching saves $17.00 on Astra's session, 62% of what the same tokens would cost uncached, and $5.40 on Kimi K3, or 65%. For Kimi K3, Moonshot puts the cache hit rate in coding workloads on its official API above 90%. The cache example on our homepage breaks the session down.
Both list a 1.05M context window. Astra accepts up to 922K input tokens and writes up to 128K, and requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Kimi K3's output defaults to 131,072 tokens per request and can be raised to the full window.
Where each runs and how OpenAI and Moonshot pitch them
GPT-6 Astra is the default model in Codex CLI's bundled model list, version 0.158.0, where it starts at low reasoning effort. Astra is available in the Codex app, CLI, and IDE extension but not in Codex cloud, and Codex is built around OpenAI's models. OpenCode, OpenRouter, and GitHub Copilot offer Astra as well.
Kimi K3 is offered in Cursor, OpenCode, OpenRouter, and GitHub Copilot, and on Moonshot's own API after a minimum $1 top-up. Kimi K3 ships open weights under Moonshot's own license, so self-hosting is possible, and this page does not estimate what that costs.
OpenAI calls Astra "our most capable model, built for the hardest end-to-end work" and says it follows instructions better than previous models. Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters." Each is its maker's top model, and neither claim measures your code. EveryToken prices Astra at OpenAI's rates from Codex and OpenCode history, and Kimi K3 from OpenRouter's catalog when you use it through OpenRouter.
Prompt caching
How each maker bills cached tokens
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what Kimi K3 and GPT-6 Astra really cost you.
everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much more does GPT-6 Astra cost than Kimi K3?
Astra's input, output, and cache-hit rates are 3.3x Kimi K3's. On the example agentic session it costs $10.50 against $2.85, or 3.7x, because OpenAI's cache writes carry a 1.25x premium and Kimi's 5-minute writes do not.
Does GPT-6 Astra cost less per task than its price suggests?
OpenAI says Astra reaches stronger results with substantially fewer output tokens in several evaluations, for a lower estimated cost per task. That is OpenAI's claim about Astra, and whether it holds depends on your work. The per-token figures here do not account for it.
Which model does Codex use by default?
GPT-6 Astra is the default in Codex CLI's bundled model list, version 0.158.0. Codex, OpenAI's own coding agent, is built around and defaults to OpenAI's models.
What is the output limit on each model?
GPT-6 Astra writes up to 128K tokens per request. Kimi K3 defaults to 131,072 and can be raised to its full 1.05M context window.
Sources
- Kimi API: Chat pricing
- Kimi API: Kimi K3 quickstart
- Kimi: Kimi K3
- OpenRouter: Kimi K3
- OpenCode docs: Zen
- Cursor docs: Models
- GitHub Docs: Supported AI models in Copilot
- OpenAI: API pricing
- OpenAI docs: GPT-6 Astra
- OpenAI: Using the latest model
- OpenAI: API changelog
- Codex docs: Models
- Codex CLI: bundled model catalog
- OpenRouter: GPT-6 Astra
- OpenCode docs: Zen
- OpenAI: Prompt caching