Skip to content

Model comparison

Gemini 3.7 Flash vs Gemini 3.6 Flash for coding agents

Gemini 3.7 Flash and Gemini 3.6 Flash cost the same and are both previous-generation. How Google pitched each, and the GitHub Copilot date to know.

· Prices as of September 28, 2026

  • Gemini 3.7 Flash

    Google · Released August 13, 2026 · Previous generation

    Launched as Google's coding and agent Flash, now the previous generation that Google keeps fully supported.

    Gemini 3.7 Flash facts and comparisons
  • Gemini 3.6 Flash

    Google · Released July 21, 2026 · Previous generation

    A previous-generation, token-efficient Flash for general agentic and everyday work, launched to curb Gemini 3.5 Flash's verbosity.

    Gemini 3.6 Flash facts and comparisons

The short answer

Gemini 3.7 Flash and Gemini 3.6 Flash have the same introductory rates, so a cache-heavy agentic coding session costs $0.71 on either. Google pitched 3.6 Flash as a token-efficient Flash for general agentic and everyday work and 3.7 Flash for complex coding and multi-step execution. Both are previous-generation, neither is in Gemini CLI or Cursor, and GitHub Copilot plans to retire 3.6 Flash on October 2, 2026, in favor of Gemini 3.8 Flash.

Choose Gemini 3.7 Flash if

  • Your work is complex coding and agentic workflows, which Google names as Gemini 3.7 Flash's focus.
  • You debug and fix issues, where Google claims higher first-pass code accuracy for 3.7 Flash.
  • You use GitHub Copilot, where Gemini 3.6 Flash is scheduled to go on October 2, 2026.

Choose Gemini 3.6 Flash if

  • You are moving off Gemini 3.5 Flash and want shorter responses, since Google launched Gemini 3.6 Flash to curb 3.5 Flash's verbosity.
  • You want computer use as a built-in tool, which Google lists for 3.6 Flash.
  • Your tasks are general agentic and everyday work that mixes speed with multimodal input, the balance Google describes for 3.6 Flash.

Side by side

Specs and prices

FactGemini 3.7 FlashGemini 3.6 Flash
MakerGoogleGoogle
API model idgemini-3.7-flashgemini-3.6-flash
ReleasedAugust 13, 2026July 21, 2026
StatusPrevious generationPrevious generation
Context window1.05M tokens1.05M tokens
Max output65.5K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.75$0.75
Cache hit, per 1M$0.075$0.075
Cache write, per 1M$0.75 (same as input)$0.75 (same as input)
Output, per 1M tokens$3.75$3.75
Runs inOpenCode, OpenRouter, and GitHub CopilotOpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.7 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. Gemini 3.6 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGemini 3.7 FlashGemini 3.6 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.71$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.15$0.15
Output-heavy generation, 30K input, 80K output$0.32$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$78.38$78.38
Where the session’s cost goes
Cache writes$0.30$0.30
Cache reads$0.15$0.15
Uncached input$0.08$0.08
Output$0.19$0.19
caching saves on the session with Gemini 3.7 Flash (66%)
$1.35
caching saves on the session with Gemini 3.6 Flash (66%)
$1.35

Two previous-generation Flash models at one price

Google prices Gemini 3.7 Flash and Gemini 3.6 Flash identically: $0.75 for each million input tokens, $0.075 for each million read from the cache, and $3.75 for each million of output. Every example workload matches: $0.71 for the session, $0.15 for the uncached review, $0.32 for the output-heavy generation, and $78.38 for 110 sessions a month.

Both rates are introductory through December 31, 2026. In 2027 both will cost $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens. Google's newer Gemini 3.8 Flash sits on the same introductory rates, which matters for anyone weighing a move.

Since the rates are equal, output length is what can separate them on cost. Output is 26% of the example session. Google launched 3.6 Flash to write less than Gemini 3.5 Flash, but the sources here hold no such comparison against 3.7 Flash.

How Google pitched Gemini 3.6 Flash and Gemini 3.7 Flash

Gemini 3.6 Flash came first, on July 21, 2026, launched to curb Gemini 3.5 Flash's verbosity. Google claimed coding precision with fewer unwanted edits and added computer use as a built-in tool. It now describes 3.6 Flash as its previous-generation Flash, balancing speed and multimodal capabilities across general agentic and everyday tasks.

Gemini 3.7 Flash followed on August 13, 2026, as Google's coding and agent Flash. Its launch claims were higher first-pass code accuracy when debugging and resolving issues, and more functional layouts when generating web UIs. Google's current description of 3.7 Flash is a previous-generation model for complex coding, agentic workflows, and reliable multi-step execution, still fully supported.

Read side by side, the pitches divide the work: 3.6 Flash for broad everyday agent tasks with less verbose output, and 3.7 Flash for coding-heavy agents. Both are Google's descriptions, and both models now sit behind Gemini 3.8 Flash in Google's lineup.

Where you can still run them, and the Copilot deadline

Neither model is in Gemini CLI's model list or in Cursor. You can still reach both through OpenRouter, OpenCode, and GitHub Copilot. In Gemini CLI, the Flash in the default auto model is Gemini 3.8 Flash for API key and Vertex AI accounts.

GitHub Copilot plans to retire Gemini 3.6 Flash on October 2, 2026, and suggests Gemini 3.8 Flash. The sources here list no retirement date for 3.7 Flash. Their limits match too, a 1.05M context window and up to 65.5K output tokens, and neither adds a charge for writing the cache.

Explicit caching adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models, on top of the 10% hit price. EveryToken reads your OpenCode history on your Mac and prices each request at Google's rates, and prices requests made through OpenRouter from OpenRouter's catalog, so a move between these models, or to 3.8 Flash, is easy to compare.

Prompt caching

How Google bills cached tokens

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Gemini 3.7 Flash and Gemini 3.6 Flash really cost you.

everyaitoken reads your OpenCode and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.6 Flash being retired?

GitHub Copilot plans to retire it on October 2, 2026, and suggests Gemini 3.8 Flash. The sources here list no shutdown date on the Gemini API itself.

Do Gemini 3.7 Flash and Gemini 3.6 Flash cost the same?

Yes, to the fraction of a cent: input at $0.75, cache hits at $0.075, and output at $3.75 per million tokens, until December 31, 2026. From January 1, 2027, both move to $1.50, $0.15, and $7.50.

Can I use Gemini 3.7 Flash or Gemini 3.6 Flash in Gemini CLI?

Neither is in Gemini CLI's model list. Gemini CLI uses Gemini 3.8 Flash as the Flash half of its default auto model for Gemini API key and Vertex AI users.

Which is newer, Gemini 3.7 Flash or Gemini 3.6 Flash?

Gemini 3.7 Flash, released on August 13, 2026. Gemini 3.6 Flash came out on July 21, 2026, and Gemini 3.8 Flash on September 2, 2026.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • Gemini 2.5 Pro vs Gemini 2.5 Flash

    Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash, and since September 18, 2026, only earlier users can reach either. What each costs and where to go next.

  • Gemini 3.1 Pro Preview vs Gemini 2.5 Pro

    Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.