Model comparison
Grok Build 0.1 vs Gemini 3.8 Flash: prices now and in 2027
Gemini 3.8 Flash costs $0.71 per cached coding session against $1.00 on Grok Build 0.1, until its introductory rates end on December 31, 2026.
· Prices as of September 28, 2026
Grok Build 0.1
xAI · Released May 2026 · Preview
xAI's dedicated agentic coding model, which also answers to the older grok-code-fast ids.
Grok Build 0.1 facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
Gemini 3.8 Flash costs less on the example agentic coding session, $0.71 against $1.00 on Grok Build 0.1, mostly because its cache hits cost $0.075 per million against $0.20. Grok Build 0.1 costs less on output-heavy work, $0.19 against $0.32, since its output rate is $2 against $3.75. Flash's rates are introductory through December 31, 2026, and double from January 1, 2027, which would put every Flash rate except cached reads above Grok Build 0.1's.
Choose Grok Build 0.1 if
- Your work is output-heavy: at $2 per million output tokens against $3.75, the example generation costs $0.19 against $0.32.
- You are budgeting past 2026, when Gemini 3.8 Flash's input and output rates rise to $1.50 and $7.50.
- You want xAI's dedicated agentic coding model and reach it through OpenRouter or OpenCode.
Choose Gemini 3.8 Flash if
- Your sessions reread a cached context, where Flash's hits cost $0.075 per million against $0.20.
- You use Gemini CLI, where Gemini 3.8 Flash is the Flash half of the default auto model, or Cursor or GitHub Copilot.
- You need more than 256K tokens of context: Gemini 3.8 Flash accepts 1.05M.
- You would like to trial it at no cost, since Flash's input, output, and caching all fall under the Gemini API free tier.
Side by side
Specs and prices
| Fact | Grok Build 0.1 | Gemini 3.8 Flash |
|---|---|---|
| Maker | xAI | |
| API model id | grok-build-0.1 | gemini-3.8-flash |
| Released | May 2026 | September 2, 2026 |
| Status | Preview | Current |
| Context window | 256K tokens | 1.05M tokens |
| Max output | Not published | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $1 | $0.75 |
| Cache hit, per 1M | $0.20 | $0.075 |
| Cache write, per 1M | $1 (same as input) | $0.75 (same as input) |
| Output, per 1M tokens | $2 | $3.75 |
| Runs in | OpenCode and OpenRouter | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Grok Build 0.1: Once a prompt reaches 200K tokens, every token in the request costs $2 input, $0.40 cached, and $4 output per million. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Grok Build 0.1 | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.00 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.17 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.19 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $110.00 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.40 | $0.30 |
| Cache reads | $0.40 | $0.15 |
| Uncached input | $0.10 | $0.08 |
| Output | $0.10 | $0.19 |
- caching saves on the session with Grok Build 0.1 (62%)
- $1.60
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Gemini 3.8 Flash's introductory rates end on December 31, 2026
Gemini 3.8 Flash lists $0.75 per million input tokens, $0.075 per cache hit, and $3.75 per million output tokens. Google labels these introductory rates through December 31, 2026. From January 1, 2027 they become $1.50 input, $0.15 cached, and $7.50 output, 2x across the board.
Grok Build 0.1 lists $1 input, $0.20 cached, and $2 output, with no promotion noted. Today Flash is cheaper on input and cache hits and dearer on output. At the 2027 rates Flash's input would cost more than Grok Build 0.1's and its output nearly four times as much, while its cache hits would still cost less, $0.15 against $0.20.
The cost table uses the introductory rates. For work you expect to run into 2027, rerun the comparison at Flash's later prices before committing to it.
How the session and output-heavy results split
In the example agentic session, 2M tokens come from the cache. They cost $0.40 on Grok Build 0.1 and $0.15 on Flash, a $0.25 difference that accounts for most of the session gap. Neither maker charges a separate cache-write fee, so the 400K written tokens cost ordinary input: $0.40 on Grok Build 0.1 and $0.30 on Flash.
The session ends at $1.00 against $0.71, or $110.00 against $78.38 over 110 sessions a month. The large one-off review is closer, $0.17 against $0.15. The output-heavy generation reverses the order: $0.19 on Grok Build 0.1 against $0.32 on Flash, 41% less, because output is the one rate where Grok Build 0.1 undercuts Flash today.
Output also depends on settings the table can't show. Gemini 3.8 Flash defaults to a medium thinking level, and Gemini CLI sends high. xAI lists reasoning among Grok Build 0.1's features. On either model, the settings a tool sends change how many tokens a task produces.
Context, limits, and where each model runs
Grok Build 0.1 accepts 256K tokens, and once a prompt reaches 200K, every token in the request costs $2 input, $0.40 cached, and $4 output per million. Flash accepts 1.05M, writes up to 65.5K tokens per request, and its price notes list no long-context tier. xAI does not publish Grok Build 0.1's maximum output.
Flash runs in Gemini CLI, Google's own coding agent, as the Flash half of the default auto model for Gemini API key and Vertex AI users, and in Cursor, OpenCode, OpenRouter, and GitHub Copilot. Grok Build 0.1 runs through OpenRouter and OpenCode. xAI announced it as early access in May 2026, and it also answers to the older grok-code-fast ids.
Google's caching comes in two forms. Implicit caching is on by default, but a hit is not assured. Explicit caching assures the discount and adds storage at $0.50 to $1 per million tokens per hour on Flash models. xAI's caching is automatic, and reusing the same conversation id raises the hit rate.
What xAI and Google say about each model
Google calls Gemini 3.8 Flash "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows," and claims better robustness against prompt injection. xAI calls Grok Build 0.1 "xAI's coding model, trained specifically for agentic coding workflows." Both makers point these models at agentic coding.
EveryToken prices Gemini 3.8 Flash at Google's rates from Gemini CLI, Cursor, or OpenCode history, and prices Grok Build 0.1 through OpenRouter, from OpenRouter's catalog. This page's tables price it at xAI's own API rates instead.
Prompt caching
How each maker bills cached tokens
xAI
The xAI API caches repeated prompt prefixes automatically. Sending the same conversation id with each request raises the hit rate.
xAI lists no fee for writing the cache. A cache hit costs $0.50 per million tokens on Grok 4.7 and $0.20 on Grok Build 0.1.
Source: xAI docs: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Grok Build 0.1 and Gemini 3.8 Flash really cost you.
everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Will Gemini 3.8 Flash get more expensive?
Google lists its current rates as introductory through December 31, 2026. From January 1, 2027, Flash costs $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Which is cheaper for a coding session, Grok Build 0.1 or Gemini 3.8 Flash?
At today's rates, Gemini 3.8 Flash: the example agentic session costs $0.71 against $1.00. For output-heavy work Grok Build 0.1 costs less, $0.19 against $0.32.
How do the context windows of Grok Build 0.1 and Gemini 3.8 Flash compare?
Gemini 3.8 Flash, at 1.05M tokens against 256K on Grok Build 0.1. Grok Build 0.1 also doubles every rate once a prompt reaches 200K tokens.
Is there a free way to try either model?
The Gemini API free tier covers Gemini 3.8 Flash's input, output, and caching. No free tier is listed here for Grok Build 0.1.
Sources
- xAI docs: Grok Build 0.1
- xAI docs: Release notes
- OpenRouter: Grok Build 0.1
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- xAI docs: Prompt caching
- Google: Context caching