Model comparison
Grok 4.7 vs Claude Sonnet 5: a near tie on cached sessions
Grok 4.7 and Claude Sonnet 5 both charge $2 per million input tokens. A cached coding session costs $2.30 against $2.40, but output-heavy work splits wider.
· Prices as of September 28, 2026
Grok 4.7
xAI · Released September 21, 2026
xAI's top model for coding and knowledge work, which xAI says works longer on hard tasks and checks its own work more carefully.
Grok 4.7 facts and comparisonsClaude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisons
The short answer
Grok 4.7 and Claude Sonnet 5 both list $2 per million input tokens, and the example agentic coding session costs nearly the same on each: $2.30 on Grok 4.7 and $2.40 on Sonnet 5. The split widens on output-heavy work, where Grok 4.7's $6 output rate against $10 makes it 37% cheaper, while Sonnet 5's $0.20 cache hits favor long sessions that reread their context. Sonnet 5 also brings a 1M window and Claude Code, and Grok 4.7 stops at 500K with higher rates once a prompt reaches 200K tokens.
Choose Grok 4.7 if
- Your work writes a lot of output: at $6 per million against $10, the example generation costs $0.54 against $0.86.
- You want cache writes without a premium, since xAI lists no fee for writing the cache and bills written tokens as input.
- You code in Cursor, where Grok 4.7 is listed as trained jointly by Cursor and xAI.
- You are drawn to xAI's description of a model that works longer on difficult tasks and checks its own work more carefully.
Choose Claude Sonnet 5 if
- Your sessions reread a large context many times, since a Sonnet 5 cache hit costs $0.20 per million against $0.50.
- Your team lives in Claude Code, where typing the sonnet alias gets you Sonnet 5 on the Anthropic API.
- You need more than 500K tokens of context, or long prompts without a price step at 200K.
- You want 128K of output per request, a limit xAI does not publish for Grok 4.7.
Side by side
Specs and prices
| Fact | Grok 4.7 | Claude Sonnet 5 |
|---|---|---|
| Maker | xAI | Anthropic |
| API model id | grok-4.7 | claude-sonnet-5 |
| Released | September 21, 2026 | June 30, 2026 |
| Status | Current | Current |
| Context window | 500K tokens | 1M tokens |
| Max output | Not published | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $2 |
| Cache hit, per 1M | $0.50 | $0.20 |
| Cache write, per 1M | $2 (same as input) | $2.50 (5-minute), $4 (1-hour) |
| Output, per 1M tokens | $6 | $10 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Grok 4.7: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Grok 4.7: Once a prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. The US regional endpoint costs 10% more. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Grok 4.7 | Claude Sonnet 5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.30 | $2.40 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $253.00 | $264.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $1.30 |
| Cache reads | $1.00 | $0.40 |
| Uncached input | $0.20 | $0.20 |
| Output | $0.30 | $0.50 |
- caching saves on the session with Grok 4.7 (57%)
- $3.00
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
Same $2 input price, different cost structures
Grok 4.7 and Claude Sonnet 5 both charge $2 per million input tokens, so fresh input costs $0.20 in the example session on either. After that the rate cards diverge in three places, output, cache writes, and cache hits, and they do not all lean the same way.
Output favors Grok 4.7, at $6 per million against $10, 40% less. Cache writes favor it too: xAI lists no write fee, so the 400K tokens written to the cache cost $0.80, while Anthropic's mix of 5-minute writes at $2.50 and 1-hour writes at $4 comes to $1.30. Cache hits favor Sonnet 5. A hit costs $0.20 per million there, 10% of input, against $0.50 on Grok 4.7, 25% of input, so the session's 2M cached tokens cost $1.00 on Grok 4.7 and $0.40 on Sonnet 5.
Add the lines and the session costs $2.30 on Grok 4.7 against $2.40 on Sonnet 5, a gap of $0.10, or $11.00 over 110 sessions a month. Grok 4.7 saves $0.50 on writes and $0.20 on output, then gives back $0.60 on reads. Uncached work has no reads to give back, so the large one-off review costs $0.36 against $0.40 and the output-heavy generation $0.54 against $0.86.
Which model is cheaper for long agentic sessions?
The more often a session rereads its cached context for each token it writes, the more Sonnet 5's cheaper hits count. The example reads 2M cached tokens against 400K written. A session that runs longer on a stable context tilts toward Sonnet 5, and one that rewrites its context often, or produces long answers, tilts toward Grok 4.7.
How Claude Code is set up matters here. Its main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key, and the table assumes an even split of the two. If every write were a 5-minute write at $2.50, Sonnet 5 would cost less than Grok 4.7 on this session. Subscriptions are priced differently from API rates altogether, so the table is an API-equivalent estimate either way.
Output length depends on settings the table can't show. Sonnet 5 runs adaptive thinking by default at high effort, and xAI describes Grok 4.7 as working longer on difficult tasks. Either can write more than the example's 50K tokens, and each extra output token widens Grok 4.7's price advantage.
Context limits: 500K with a 200K step against 1M
Sonnet 5 accepts 1M tokens of context and writes up to 128K of output. Grok 4.7 accepts 500K, and xAI does not publish its maximum output. For most edits and reviews neither limit comes into play.
The 200K line matters more than the window. Once a Grok 4.7 prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. Past that point every Grok 4.7 rate is higher than Sonnet 5's: $4 against $2 for input, $1 against $0.20 for hits, and $12 against $10 for output. The price data here lists no long-context tier for Sonnet 5.
xAI's US regional endpoint also adds 10% to Grok 4.7's rates. The tables use xAI's and Anthropic's own API prices, not what any subscription or third-party tool charges.
How xAI and Anthropic position the two models
xAI calls Grok 4.7 "our most capable model for coding and knowledge work" and says it rests on a new, larger base model trained with a longer reinforcement-learning run on many-hour tasks. Anthropic calls Sonnet 5 its balance of speed and intelligence, says it is "built to be the most agentic Sonnet model yet," and claims performance close to Claude Opus 4.8 at lower prices. Both run in Cursor, OpenCode, OpenRouter, and GitHub Copilot, and Sonnet 5 also runs in Claude Code, Anthropic's own coding agent.
Tokenizers differ between the makers, and Sonnet 5's counts about 1.0 to 1.35x as many tokens as Claude Sonnet 4.6 for the same text. To compare the two on your own sessions, EveryToken prices Sonnet 5 at Anthropic's rates from Claude Code, Cursor, or OpenCode history, and Grok 4.7 when you use it through OpenRouter, from OpenRouter's catalog.
Prompt caching
How each maker bills cached tokens
xAI
The xAI API caches repeated prompt prefixes automatically. Sending the same conversation id with each request raises the hit rate.
xAI lists no fee for writing the cache. A cache hit costs $0.50 per million tokens on Grok 4.7 and $0.20 on Grok Build 0.1.
Source: xAI docs: Prompt caching
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what Grok 4.7 and Claude Sonnet 5 really cost you.
everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Grok 4.7 cheaper than Claude Sonnet 5?
On uncached and output-heavy work, yes: the example generation costs $0.54 against $0.86. On the cached agentic session the two nearly tie, $2.30 against $2.40, because Sonnet 5's cache hits cost $0.20 per million against $0.50. Which one costs less depends on how much of your work is cache reads.
Does Grok 4.7 charge for cache writes?
xAI lists no fee for writing the cache, so written tokens cost the $2 input rate. Sonnet 5 charges $2.50 per million for a 5-minute write and $4 for a 1-hour write.
How large is the context window on each model?
Grok 4.7 accepts 500K tokens and Claude Sonnet 5 accepts 1M. Grok 4.7's rates double once a prompt reaches 200K tokens, so long prompts cost more there.
Where can I use both models?
Cursor, OpenCode, OpenRouter, and GitHub Copilot offer both. Claude Code offers Sonnet 5, and its sonnet alias resolves to it on the Anthropic API.
Sources
- xAI docs: Grok 4.7
- xAI: Grok 4.7
- OpenRouter: Grok 4.7
- OpenCode docs: Zen
- Cursor docs: Models
- GitHub Docs: Supported AI models in Copilot
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- xAI docs: Prompt caching
- Anthropic: Prompt caching