Skip to content

Model comparison

GLM-5.3-Flash vs Gemini 3.8 Flash: prices now and in 2027

GLM-5.3-Flash costs $0.16 for a cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end December 31, 2026. What else differs.

· Prices as of September 28, 2026

  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

GLM-5.3-Flash costs less on every line, $0.16 against $0.71 for the example agentic coding session, while Gemini 3.8 Flash costs 8x as much on output-heavy work, where it charges $3.75 per million output tokens against $0.50. Gemini 3.8 Flash is the choice inside Gemini CLI, Cursor, or GitHub Copilot, and its free tier covers it, but its current prices are introductory rates that end on December 31, 2026. GLM-5.3-Flash suits cost-driven work through OpenRouter or OpenCode, with MIT-licensed weights and 128K of output per response.

Choose GLM-5.3-Flash if

  • Output is a large share of your work, and GLM-5.3-Flash bills it at $0.50 per million against $3.75.
  • You are budgeting past 2026, when Gemini 3.8 Flash moves to $1.50 input and $7.50 output on January 1, 2027.
  • Single responses need to run past 65.5K tokens; GLM-5.3-Flash allows 128K.
  • You want weights you can run yourself, published under the MIT license.

Choose Gemini 3.8 Flash if

  • Gemini CLI is your agent, and with a Gemini API key or Vertex AI its default auto model uses Gemini 3.8 Flash for Flash requests.
  • You want to begin at no cost, since the Gemini API free tier covers Gemini 3.8 Flash's input, output, and caching.
  • Your editor is Cursor or your assistant is GitHub Copilot, and both list Gemini 3.8 Flash but not GLM-5.3-Flash.
  • You want a model Google aims at long-horizon software engineering, multi-file refactoring, and deterministic tool execution.

Side by side

Specs and prices

FactGLM-5.3-FlashGemini 3.8 Flash
MakerZ.aiGoogle
API model idglm-5.3-flashgemini-3.8-flash
ReleasedAugust 26, 2026September 2, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsYesNo
Input, per 1M tokens$0.15$0.75
Cache hit, per 1M$0.03$0.075
Cache write, per 1M$0.15 (same as input)$0.75 (same as input)
Output, per 1M tokens$0.50$3.75
Runs inOpenCode and OpenRouterCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3-FlashGemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.16$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.03$0.15
Output-heavy generation, 30K input, 80K output$0.04$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$17.60$78.38
Where the session’s cost goes
Cache writes$0.06$0.30
Cache reads$0.06$0.15
Uncached input$0.02$0.08
Output$0.03$0.19
caching saves on the session with GLM-5.3-Flash (60%)
$0.24
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Why the gap narrows on the cached session

On uncached work the rate cards decide. GLM-5.3-Flash charges $0.15 input and $0.50 output per million on Z.ai's API, and Gemini 3.8 Flash charges $0.75 and $3.75. The large one-off review costs $0.03 against $0.15, 5x, and the output-heavy generation $0.04 against $0.32, 8x.

The example session is 4.4x apart, $0.16 against $0.71, a smaller gap than either uncached workload. The reason is the cache. Google bills a hit at 10% of input, $0.075 per million, while Z.ai bills $0.03, 20% of input. Gemini's steeper discount pulls its read cost closer: 2M cached tokens cost $0.15 on Gemini against $0.06 on GLM-5.3-Flash, only 2.5x apart.

Neither maker adds a premium for writing the cache. Z.ai lists no write fee and Google publishes no separate write price, so written tokens cost ordinary input on both: $0.06 on GLM-5.3-Flash and $0.30 on Gemini for the session's 400K. Caching removes 66% of the uncached session cost on Gemini and 60% on GLM-5.3-Flash.

Gemini 3.8 Flash prices after December 31, 2026

The Gemini 3.8 Flash rates in these tables are introductory, and Google lists them through December 31, 2026. On January 1, 2027 they become $1.50 input, $0.15 cached, and $7.50 output per million tokens, so the same comparison in 2027 would show a wider gap to GLM-5.3-Flash, provided Z.ai's price holds.

Google's explicit caching is another cost to plan for. Implicit caching is on by default for Gemini 2.5 and newer, but a hit is not certain. Explicit caching gives an assured discount and adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models. Z.ai caches repeated context automatically and says cached input storage is free for a limited time.

The Gemini API free tier covers Gemini 3.8 Flash's input, output, and caching, which makes it easy to try before paying. The tables use paid-tier rates.

Output limits, tools, and what each maker claims

Both accept roughly 1M tokens of context: 1M on GLM-5.3-Flash and 1.05M on Gemini 3.8 Flash. Output limits differ: 128K per response on GLM-5.3-Flash and 65.5K on Gemini. Gemini 3.8 Flash's default thinking level is medium and Gemini CLI sends high, and the thinking level changes how many output tokens a task uses.

Google's own description of Gemini 3.8 Flash is "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Google adds that it holds up better against prompt injection. Z.ai's pitch for GLM-5.3-Flash is a low-cost model with native multimodal visual coding, one that looks at interfaces and rendered results to test and improve its work, and that outperforms GLM-5.2 at a tenth of the price.

Gemini CLI is Google's own coding agent and runs Gemini models; Gemini 3.8 Flash is also in Cursor, OpenRouter, OpenCode, and GitHub Copilot. GLM-5.3-Flash runs through OpenRouter and OpenCode. On OpenRouter a request for an open-weight model goes to one of several providers, and their prices can differ from Z.ai's.

EveryToken prices Gemini 3.8 Flash at Google's rates in Gemini CLI, Cursor, and OpenCode, and prices GLM-5.3-Flash through OpenRouter from OpenRouter's catalog.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GLM-5.3-Flash and Gemini 3.8 Flash really cost you.

everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3-Flash cheaper than Gemini 3.8 Flash?

Yes, at current list prices on every line. The example cached session costs $0.16 against $0.71, and 110 sessions a month come to $17.60 against $78.38.

Will Gemini 3.8 Flash get more expensive?

Its current prices are introductory and run through December 31, 2026. Starting January 1, 2027, Google charges $1.50 input, $0.15 cached, and $7.50 output per million tokens for it.

Which model writes longer responses?

GLM-5.3-Flash lists 128K output tokens per response, and Gemini 3.8 Flash 65.5K. Their context windows are 1M and 1.05M.

Is there a free way to try either model?

Gemini 3.8 Flash can be tried on the Gemini API free tier, which includes its input, output, and caching. GLM-5.3-Flash has open weights under the MIT license, so you can run it on your own hardware, at whatever that hardware costs.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.