Model comparison
Gemini 3.7 Flash vs Gemini 3.6 Flash for coding agents
Gemini 3.7 Flash and Gemini 3.6 Flash cost the same and are both previous-generation. How Google pitched each, and the GitHub Copilot date to know.
· Prices as of September 28, 2026
Gemini 3.7 Flash
Google · Released August 13, 2026 · Previous generation
Launched as Google's coding and agent Flash, now the previous generation that Google keeps fully supported.
Gemini 3.7 Flash facts and comparisonsGemini 3.6 Flash
Google · Released July 21, 2026 · Previous generation
A previous-generation, token-efficient Flash for general agentic and everyday work, launched to curb Gemini 3.5 Flash's verbosity.
Gemini 3.6 Flash facts and comparisons
The short answer
Gemini 3.7 Flash and Gemini 3.6 Flash have the same introductory rates, so a cache-heavy agentic coding session costs $0.71 on either. Google pitched 3.6 Flash as a token-efficient Flash for general agentic and everyday work and 3.7 Flash for complex coding and multi-step execution. Both are previous-generation, neither is in Gemini CLI or Cursor, and GitHub Copilot plans to retire 3.6 Flash on October 2, 2026, in favor of Gemini 3.8 Flash.
Choose Gemini 3.7 Flash if
- Your work is complex coding and agentic workflows, which Google names as Gemini 3.7 Flash's focus.
- You debug and fix issues, where Google claims higher first-pass code accuracy for 3.7 Flash.
- You use GitHub Copilot, where Gemini 3.6 Flash is scheduled to go on October 2, 2026.
Choose Gemini 3.6 Flash if
- You are moving off Gemini 3.5 Flash and want shorter responses, since Google launched Gemini 3.6 Flash to curb 3.5 Flash's verbosity.
- You want computer use as a built-in tool, which Google lists for 3.6 Flash.
- Your tasks are general agentic and everyday work that mixes speed with multimodal input, the balance Google describes for 3.6 Flash.
Side by side
Specs and prices
| Fact | Gemini 3.7 Flash | Gemini 3.6 Flash |
|---|---|---|
| Maker | ||
| API model id | gemini-3.7-flash | gemini-3.6-flash |
| Released | August 13, 2026 | July 21, 2026 |
| Status | Previous generation | Previous generation |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.75 | $0.75 |
| Cache hit, per 1M | $0.075 | $0.075 |
| Cache write, per 1M | $0.75 (same as input) | $0.75 (same as input) |
| Output, per 1M tokens | $3.75 | $3.75 |
| Runs in | OpenCode, OpenRouter, and GitHub Copilot | OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.7 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. Gemini 3.6 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 3.7 Flash | Gemini 3.6 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.71 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.15 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.32 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $78.38 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.30 | $0.30 |
| Cache reads | $0.15 | $0.15 |
| Uncached input | $0.08 | $0.08 |
| Output | $0.19 | $0.19 |
- caching saves on the session with Gemini 3.7 Flash (66%)
- $1.35
- caching saves on the session with Gemini 3.6 Flash (66%)
- $1.35
Two previous-generation Flash models at one price
Google prices Gemini 3.7 Flash and Gemini 3.6 Flash identically: $0.75 for each million input tokens, $0.075 for each million read from the cache, and $3.75 for each million of output. Every example workload matches: $0.71 for the session, $0.15 for the uncached review, $0.32 for the output-heavy generation, and $78.38 for 110 sessions a month.
Both rates are introductory through December 31, 2026. In 2027 both will cost $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens. Google's newer Gemini 3.8 Flash sits on the same introductory rates, which matters for anyone weighing a move.
Since the rates are equal, output length is what can separate them on cost. Output is 26% of the example session. Google launched 3.6 Flash to write less than Gemini 3.5 Flash, but the sources here hold no such comparison against 3.7 Flash.
How Google pitched Gemini 3.6 Flash and Gemini 3.7 Flash
Gemini 3.6 Flash came first, on July 21, 2026, launched to curb Gemini 3.5 Flash's verbosity. Google claimed coding precision with fewer unwanted edits and added computer use as a built-in tool. It now describes 3.6 Flash as its previous-generation Flash, balancing speed and multimodal capabilities across general agentic and everyday tasks.
Gemini 3.7 Flash followed on August 13, 2026, as Google's coding and agent Flash. Its launch claims were higher first-pass code accuracy when debugging and resolving issues, and more functional layouts when generating web UIs. Google's current description of 3.7 Flash is a previous-generation model for complex coding, agentic workflows, and reliable multi-step execution, still fully supported.
Read side by side, the pitches divide the work: 3.6 Flash for broad everyday agent tasks with less verbose output, and 3.7 Flash for coding-heavy agents. Both are Google's descriptions, and both models now sit behind Gemini 3.8 Flash in Google's lineup.
Where you can still run them, and the Copilot deadline
Neither model is in Gemini CLI's model list or in Cursor. You can still reach both through OpenRouter, OpenCode, and GitHub Copilot. In Gemini CLI, the Flash in the default auto model is Gemini 3.8 Flash for API key and Vertex AI accounts.
GitHub Copilot plans to retire Gemini 3.6 Flash on October 2, 2026, and suggests Gemini 3.8 Flash. The sources here list no retirement date for 3.7 Flash. Their limits match too, a 1.05M context window and up to 65.5K output tokens, and neither adds a charge for writing the cache.
Explicit caching adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models, on top of the 10% hit price. EveryToken reads your OpenCode history on your Mac and prices each request at Google's rates, and prices requests made through OpenRouter from OpenRouter's catalog, so a move between these models, or to 3.8 Flash, is easy to compare.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 3.7 Flash and Gemini 3.6 Flash really cost you.
everyaitoken reads your OpenCode and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.6 Flash being retired?
GitHub Copilot plans to retire it on October 2, 2026, and suggests Gemini 3.8 Flash. The sources here list no shutdown date on the Gemini API itself.
Do Gemini 3.7 Flash and Gemini 3.6 Flash cost the same?
Yes, to the fraction of a cent: input at $0.75, cache hits at $0.075, and output at $3.75 per million tokens, until December 31, 2026. From January 1, 2027, both move to $1.50, $0.15, and $7.50.
Can I use Gemini 3.7 Flash or Gemini 3.6 Flash in Gemini CLI?
Neither is in Gemini CLI's model list. Gemini CLI uses Gemini 3.8 Flash as the Flash half of its default auto model for Gemini API key and Vertex AI users.
Which is newer, Gemini 3.7 Flash or Gemini 3.6 Flash?
Gemini 3.7 Flash, released on August 13, 2026. Gemini 3.6 Flash came out on July 21, 2026, and Gemini 3.8 Flash on September 2, 2026.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.7 Flash
- Google: Gemini models
- Google: Introducing Gemini 3.7 Flash
- Gemini CLI source: model configuration
- Cursor docs: Models
- OpenRouter: Gemini 3.7 Flash
- Google docs: Gemini 3.6 Flash
- Google: Gemini 3.6 Flash and 3.5 Flash-Lite
- OpenRouter: Gemini 3.6 Flash
- GitHub Docs: Supported AI models in Copilot
- Google: Context caching