Model comparison
Grok 4.7 vs Kimi K3: $6 or $15 per million output tokens
Grok 4.7 and Kimi K3 both run in Cursor and GitHub Copilot. Grok 4.7 costs less on every workload, while Kimi K3 adds open weights and a 1.05M window.
· Prices as of September 28, 2026
Grok 4.7
xAI · Released September 21, 2026
xAI's top model for coding and knowledge work, which xAI says works longer on hard tasks and checks its own work more carefully.
Grok 4.7 facts and comparisonsKimi K3
Moonshot AI · Released July 16, 2026
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Kimi K3 facts and comparisons
The short answer
Grok 4.7 costs less than Kimi K3 on every example workload, but by very different margins: $2.30 against $2.85 on the cached agentic coding session, and $0.54 against $1.29 on output-heavy work, since output costs $6 per million against $15. Kimi K3's cheaper cache hits, $0.30 against $0.50, are what keep the session close. Kimi K3 adds open weights, a 1.05M context window, and up to 1.05M tokens of output, while Grok 4.7 stops at 500K and doubles its rates once a prompt reaches 200K tokens.
Choose Grok 4.7 if
- Your work is output-heavy: Grok 4.7 bills output at $6 per million against $15 on Kimi K3.
- You want the lower total on every example workload, from $0.36 against $0.60 on the one-off review to $253.00 against $313.50 for a month of sessions.
- You pick models in Cursor, whose listing describes Grok 4.7 as a joint effort of Cursor and xAI and adds a faster Grok 4.7 Fast option.
Choose Kimi K3 if
- Downloadable weights matter to you, and Moonshot releases Kimi K3's under a license of its own.
- You need more than 500K tokens of context or very long responses: Kimi K3 accepts 1.05M and can raise its output to the full window.
- Your prompts often pass 200K tokens, where Grok 4.7 moves to $4 input and $12 output, while Kimi K3's price notes show no long-context tier.
- Your sessions reread a cached context much more than the example does: a Kimi K3 hit costs $0.30 against $0.50, and Moonshot reports a cache hit rate above 90% in coding workloads on its API.
Side by side
Specs and prices
| Fact | Grok 4.7 | Kimi K3 |
|---|---|---|
| Maker | xAI | Moonshot AI |
| API model id | grok-4.7 | kimi-k3 |
| Released | September 21, 2026 | July 16, 2026 |
| Status | Current | Current |
| Context window | 500K tokens | 1.05M tokens |
| Max output | Not published | 1.05M tokens |
| Open weights | No | Yes |
| Input, per 1M tokens | $2 | $3 |
| Cache hit, per 1M | $0.50 | $0.30 |
| Cache write, per 1M | $2 (same as input) | $3 (same as input) |
| Output, per 1M tokens | $6 | $15 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Grok 4.7: Once a prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. The US regional endpoint costs 10% more. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Grok 4.7 | Kimi K3 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.30 | $2.85 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.60 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $1.29 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $253.00 | $313.50 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $1.20 |
| Cache reads | $1.00 | $0.60 |
| Uncached input | $0.20 | $0.30 |
| Output | $0.30 | $0.75 |
- caching saves on the session with Grok 4.7 (57%)
- $3.00
- caching saves on the session with Kimi K3 (65%)
- $5.40
Grok 4.7 and Kimi K3 run in the same four tools
Grok 4.7 and Kimi K3 are both offered in Cursor, GitHub Copilot, OpenCode, and OpenRouter, the multi-maker tools compared here. Choosing between them rarely means switching tools, so the rate cards and limits carry more of the decision.
xAI calls Grok 4.7 "our most capable model for coding and knowledge work" and says it works longer on difficult tasks and checks its own work more carefully. Cursor describes it as trained jointly by Cursor and xAI. Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," and names long engineering tasks with minimal supervision, large codebases, and terminal tools as strengths, along with visual reasoning from screenshots for frontend and game work.
Output is where the gap opens
Grok 4.7 charges $2 per million input tokens and $6 per million output tokens. Kimi K3 charges $3 and $15. Output is 2.5x dearer on Kimi K3 and input 1.5x, which is why the output-heavy generation shows a 2.4x gap, $0.54 against $1.29, while the large one-off review, mostly input, shows 1.7x, $0.36 against $0.60.
Caching pulls the two closer. A Kimi K3 cache hit costs $0.30, one-tenth of its input price, and a Grok 4.7 hit $0.50, 25% of input. The example session's 2M cached tokens cost $0.60 on Kimi K3 and $1.00 on Grok 4.7. Neither adds a premium for the default write: Kimi bills its 5-minute write at $3, the same as input, and xAI lists no write fee at all. The session ends at $2.30 against $2.85, only 1.2x apart, or $253.00 against $313.50 over 110 sessions.
Kimi also sells a 1-hour cache, with writes at $6 per million, and each hit restarts the cache lifetime at no charge. The table prices Kimi's writes at the 5-minute default, so a setup that asks for the 1-hour cache would pay more for writes than shown. At the default, caching saves 65% of Kimi K3's uncached session cost.
Context, output, and long-prompt pricing
Kimi K3 accepts 1.05M tokens of context and can write up to the same 1.05M of output, though output defaults to 131,072 tokens per request until you raise it. Grok 4.7 stops at 500K of context, and xAI publishes no output ceiling for it.
Grok 4.7 also has a price step. Once a prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. Kimi K3's price notes list no long-context tier. Above 200K, then, Grok 4.7's input and cached rates pass Kimi K3's, while its output, at $12 against $15, stays lower.
Two access details round this out. xAI's US regional endpoint adds 10% to Grok 4.7's rates, and on the Kimi API, access unlocks after a minimum $1 top-up.
Open weights and pricing your own usage
Moonshot publishes Kimi K3's weights under its own Kimi K3 license, and Grok 4.7's weights are not published. On OpenRouter, a request for an open-weight model like Kimi K3 is served by one of several providers, whose prices can differ from Moonshot's own API. The tables use each maker's own API price, and this post does not estimate self-hosting costs.
EveryToken prices both models when you use them through OpenRouter, from OpenRouter's catalog. For Kimi K3 that catalog price can differ from the list price in the table, which is one more reason to check real usage rather than a rate card.
Prompt caching
How each maker bills cached tokens
xAI
The xAI API caches repeated prompt prefixes automatically. Sending the same conversation id with each request raises the hit rate.
xAI lists no fee for writing the cache. A cache hit costs $0.50 per million tokens on Grok 4.7 and $0.20 on Grok Build 0.1.
Source: xAI docs: Prompt caching
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
Your own numbers
See what Grok 4.7 and Kimi K3 really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Grok 4.7 cheaper than Kimi K3?
Yes, on every example workload. The gap is small on the cached agentic session, $2.30 against $2.85, and large on output-heavy work, $0.54 against $1.29, because Kimi K3's output costs $15 per million against $6.
Which of the two accepts more context?
Kimi K3, with 1.05M tokens against 500K on Grok 4.7. Grok 4.7 also doubles its rates once a prompt reaches 200K tokens.
Does Kimi K3 charge for cache writes?
Kimi bills writes separately: $3 per million for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. xAI lists no write fee for Grok 4.7.
Where can I use both models?
Both run in Cursor, GitHub Copilot, OpenCode, and OpenRouter. In Cursor, Grok 4.7 also comes in a faster Grok 4.7 Fast version at double the rates.