Model comparison
Grok 4.7 vs Gemini 3.1 Pro Preview: two 200K price lines
Grok 4.7 and Gemini 3.1 Pro Preview both charge $2 input and raise rates at 200K tokens. Gemini costs less on a cached session, Grok on output-heavy work.
· Prices as of September 28, 2026
Grok 4.7
xAI · Released September 21, 2026
xAI's top model for coding and knowledge work, which xAI says works longer on hard tasks and checks its own work more carefully.
Grok 4.7 facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
Gemini 3.1 Pro Preview costs less on the example agentic coding session, $2.00 against $2.30 on Grok 4.7, because both bill cache writes as input and Gemini's cache hits cost $0.20 per million against $0.50. Grok 4.7 costs less wherever output dominates, at $6 per million against $12, so the output-heavy generation runs $0.54 against $1.02. Both raise their rates at 200K tokens, but Gemini 3.1 Pro Preview is still a preview with a 1.05M window, while Grok 4.7 is a current model with a 500K window.
Choose Grok 4.7 if
- Your work writes a lot of output, which costs $6 per million on Grok 4.7 against $12.
- You use GitHub Copilot, which offers Grok 4.7 and retired Gemini 3.1 Pro Preview on September 1, 2026.
- You want a current model rather than a preview: xAI released Grok 4.7 on September 21, 2026.
- You work in Cursor, whose listing describes Grok 4.7 as trained jointly by Cursor and xAI.
Choose Gemini 3.1 Pro Preview if
- Your sessions reread a large cached context, where a hit costs $0.20 per million against $0.50.
- You use Gemini CLI, where Gemini 3.1 Pro Preview is the Pro half of the default auto model.
- You need more than 500K tokens of context: Gemini 3.1 Pro Preview accepts 1.05M.
- Your long prompts are mostly cached, since above 200K tokens its cached rate is $0.40 against $1 on Grok 4.7.
Side by side
Specs and prices
| Fact | Grok 4.7 | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | xAI | |
| API model id | grok-4.7 | gemini-3.1-pro-preview |
| Released | September 21, 2026 | February 19, 2026 |
| Status | Current | Preview |
| Context window | 500K tokens | 1.05M tokens |
| Max output | Not published | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $2 |
| Cache hit, per 1M | $0.50 | $0.20 |
| Cache write, per 1M | $2 (same as input) | $2 (same as input) |
| Output, per 1M tokens | $6 | $12 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Grok 4.7: Once a prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. The US regional endpoint costs 10% more. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Grok 4.7 | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.30 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $253.00 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $0.80 |
| Cache reads | $1.00 | $0.40 |
| Uncached input | $0.20 | $0.20 |
| Output | $0.30 | $0.60 |
- caching saves on the session with Grok 4.7 (57%)
- $3.00
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
Both models change price at 200K tokens
Grok 4.7 and Gemini 3.1 Pro Preview share a pricing trait: each has a long-context rate that starts at 200K tokens. xAI's rule is that once a prompt reaches 200K tokens, every token in the request costs $4 input, $1 cached, and $12 output per million. Google's is that prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million.
Above the line, input is level at $4. Cached tokens cost less on Gemini 3.1 Pro Preview, $0.40 against $1, and output costs less on Grok 4.7, $12 against $18. The standard-rate pattern carries over into long prompts: Gemini is cheaper to reread, and Grok is cheaper for what the model writes.
The windows differ. Grok 4.7 accepts 500K tokens and Gemini 3.1 Pro Preview accepts 1.05M, so Gemini can take a prompt far past its long-context line. Google caps output at 65.5K tokens per request, and xAI does not publish an output limit for Grok 4.7.
Where the $0.30 session gap comes from
Below 200K both list $2 per million input tokens, and both bill cache writes as ordinary input, since neither xAI nor Google charges a separate write fee. In the example session the 400K written tokens cost $0.80 on each model, and fresh input costs $0.20 on each.
The difference sits in reads and output. The session's 2M cached tokens cost $1.00 on Grok 4.7, whose hit price is 25% of input, and $0.40 on Gemini 3.1 Pro Preview, at 10% of input. Output runs the other way: 50K tokens cost $0.30 on Grok 4.7 and $0.60 on Gemini. Net, Gemini comes to $2.00 against $2.30, a gap that grows to $33.00 over 110 sessions a month.
Uncached work favors Grok 4.7. The large one-off review costs $0.36 against $0.42, and the output-heavy generation costs $0.54 against $1.02, 47% less. Thinking cannot be turned off on Gemini 3.1 Pro Preview, and its default level is high, which bears on how much output a real task produces.
Preview status and where each model runs
Gemini 3.1 Pro Preview is Google's current Pro model and is still in preview. Google has announced Gemini 3.5 Pro, which is not yet released. In Gemini CLI, Google's own coding agent, it is the Pro half of the default auto model, and Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint at the same price, which Google says is better at prioritizing custom tools alongside bash.
Cursor, OpenCode, and OpenRouter offer both models. GitHub Copilot retired Gemini 3.1 Pro Preview on September 1, 2026, and offers Grok 4.7. Cursor describes Grok 4.7 as trained jointly by Cursor and xAI, and xAI's US regional endpoint costs 10% more than the rates in the tables.
What xAI and Google claim, and how their caches work
xAI says Grok 4.7 works longer on difficult tasks and checks its own work more carefully, and calls it "our most capable model for coding and knowledge work." Google describes Gemini 3.1 Pro Preview as offering "advanced intelligence, complex problem-solving skills, and powerful agentic and vibe coding capabilities," and points to software engineering and agentic workflows that need precise tool use.
Google's caching has a wrinkle the table leaves out. Its default, implicit caching, discounts a repeated prefix on its own, yet Google does not promise every repeat will hit. Explicit caching, which assures the discount, adds a storage charge of $4.50 per million tokens per hour on Pro models. xAI caches repeated prefixes automatically, and sending the same conversation id raises the hit rate.
EveryToken prices Gemini 3.1 Pro Preview at Google's rates from Gemini CLI, Cursor, or OpenCode history, and prices Grok 4.7 when it runs through OpenRouter, from OpenRouter's catalog. The Grok 4.7 figures in the tables come from xAI's own price list.
Prompt caching
How each maker bills cached tokens
xAI
The xAI API caches repeated prompt prefixes automatically. Sending the same conversation id with each request raises the hit rate.
xAI lists no fee for writing the cache. A cache hit costs $0.50 per million tokens on Grok 4.7 and $0.20 on Grok Build 0.1.
Source: xAI docs: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Grok 4.7 and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Grok 4.7 cheaper than Gemini 3.1 Pro Preview?
For output-heavy and uncached work, yes: the example generation costs $0.54 against $1.02. For a cache-heavy agentic session, no: Gemini 3.1 Pro Preview costs $2.00 against $2.30, because its cache hits cost $0.20 per million against $0.50.
What do both models charge above 200K tokens?
Grok 4.7 bills every token in a request that reaches 200K tokens at $4 input, $1 cached, and $12 output per million. Gemini 3.1 Pro Preview bills prompts over 200K input tokens at $4 input, $0.40 cached, and $18 output.
Is Gemini 3.1 Pro Preview still in GitHub Copilot?
No. GitHub Copilot retired it on September 1, 2026. Grok 4.7 is available there, and both models remain in Cursor, OpenCode, and OpenRouter.
Which model can write longer outputs?
Google publishes a limit of 65.5K output tokens per request for Gemini 3.1 Pro Preview. xAI does not publish a maximum output for Grok 4.7, so the two can't be compared on this point.
Sources
- xAI docs: Grok 4.7
- xAI: Grok 4.7
- OpenRouter: Grok 4.7
- OpenCode docs: Zen
- Cursor docs: Models
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- OpenCode docs: Zen
- xAI docs: Prompt caching
- Google: Context caching