Model comparison
GLM-5.3-Flash vs Gemini 3.8 Flash: prices now and in 2027
GLM-5.3-Flash costs $0.16 for a cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end December 31, 2026. What else differs.
· Prices as of September 28, 2026
GLM-5.3-Flash
Z.ai · Released August 26, 2026
Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.
GLM-5.3-Flash facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
GLM-5.3-Flash costs less on every line, $0.16 against $0.71 for the example agentic coding session, while Gemini 3.8 Flash costs 8x as much on output-heavy work, where it charges $3.75 per million output tokens against $0.50. Gemini 3.8 Flash is the choice inside Gemini CLI, Cursor, or GitHub Copilot, and its free tier covers it, but its current prices are introductory rates that end on December 31, 2026. GLM-5.3-Flash suits cost-driven work through OpenRouter or OpenCode, with MIT-licensed weights and 128K of output per response.
Choose GLM-5.3-Flash if
- Output is a large share of your work, and GLM-5.3-Flash bills it at $0.50 per million against $3.75.
- You are budgeting past 2026, when Gemini 3.8 Flash moves to $1.50 input and $7.50 output on January 1, 2027.
- Single responses need to run past 65.5K tokens; GLM-5.3-Flash allows 128K.
- You want weights you can run yourself, published under the MIT license.
Choose Gemini 3.8 Flash if
- Gemini CLI is your agent, and with a Gemini API key or Vertex AI its default auto model uses Gemini 3.8 Flash for Flash requests.
- You want to begin at no cost, since the Gemini API free tier covers Gemini 3.8 Flash's input, output, and caching.
- Your editor is Cursor or your assistant is GitHub Copilot, and both list Gemini 3.8 Flash but not GLM-5.3-Flash.
- You want a model Google aims at long-horizon software engineering, multi-file refactoring, and deterministic tool execution.
Side by side
Specs and prices
| Fact | GLM-5.3-Flash | Gemini 3.8 Flash |
|---|---|---|
| Maker | Z.ai | |
| API model id | glm-5.3-flash | gemini-3.8-flash |
| Released | August 26, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.15 | $0.75 |
| Cache hit, per 1M | $0.03 | $0.075 |
| Cache write, per 1M | $0.15 (same as input) | $0.75 (same as input) |
| Output, per 1M tokens | $0.50 | $3.75 |
| Runs in | OpenCode and OpenRouter | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GLM-5.3-Flash | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.16 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.03 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.04 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $17.60 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.06 | $0.30 |
| Cache reads | $0.06 | $0.15 |
| Uncached input | $0.02 | $0.08 |
| Output | $0.03 | $0.19 |
- caching saves on the session with GLM-5.3-Flash (60%)
- $0.24
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Why the gap narrows on the cached session
On uncached work the rate cards decide. GLM-5.3-Flash charges $0.15 input and $0.50 output per million on Z.ai's API, and Gemini 3.8 Flash charges $0.75 and $3.75. The large one-off review costs $0.03 against $0.15, 5x, and the output-heavy generation $0.04 against $0.32, 8x.
The example session is 4.4x apart, $0.16 against $0.71, a smaller gap than either uncached workload. The reason is the cache. Google bills a hit at 10% of input, $0.075 per million, while Z.ai bills $0.03, 20% of input. Gemini's steeper discount pulls its read cost closer: 2M cached tokens cost $0.15 on Gemini against $0.06 on GLM-5.3-Flash, only 2.5x apart.
Neither maker adds a premium for writing the cache. Z.ai lists no write fee and Google publishes no separate write price, so written tokens cost ordinary input on both: $0.06 on GLM-5.3-Flash and $0.30 on Gemini for the session's 400K. Caching removes 66% of the uncached session cost on Gemini and 60% on GLM-5.3-Flash.
Gemini 3.8 Flash prices after December 31, 2026
The Gemini 3.8 Flash rates in these tables are introductory, and Google lists them through December 31, 2026. On January 1, 2027 they become $1.50 input, $0.15 cached, and $7.50 output per million tokens, so the same comparison in 2027 would show a wider gap to GLM-5.3-Flash, provided Z.ai's price holds.
Google's explicit caching is another cost to plan for. Implicit caching is on by default for Gemini 2.5 and newer, but a hit is not certain. Explicit caching gives an assured discount and adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models. Z.ai caches repeated context automatically and says cached input storage is free for a limited time.
The Gemini API free tier covers Gemini 3.8 Flash's input, output, and caching, which makes it easy to try before paying. The tables use paid-tier rates.
Output limits, tools, and what each maker claims
Both accept roughly 1M tokens of context: 1M on GLM-5.3-Flash and 1.05M on Gemini 3.8 Flash. Output limits differ: 128K per response on GLM-5.3-Flash and 65.5K on Gemini. Gemini 3.8 Flash's default thinking level is medium and Gemini CLI sends high, and the thinking level changes how many output tokens a task uses.
Google's own description of Gemini 3.8 Flash is "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Google adds that it holds up better against prompt injection. Z.ai's pitch for GLM-5.3-Flash is a low-cost model with native multimodal visual coding, one that looks at interfaces and rendered results to test and improve its work, and that outperforms GLM-5.2 at a tenth of the price.
Gemini CLI is Google's own coding agent and runs Gemini models; Gemini 3.8 Flash is also in Cursor, OpenRouter, OpenCode, and GitHub Copilot. GLM-5.3-Flash runs through OpenRouter and OpenCode. On OpenRouter a request for an open-weight model goes to one of several providers, and their prices can differ from Z.ai's.
EveryToken prices Gemini 3.8 Flash at Google's rates in Gemini CLI, Cursor, and OpenCode, and prices GLM-5.3-Flash through OpenRouter from OpenRouter's catalog.
Prompt caching
How each maker bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what GLM-5.3-Flash and Gemini 3.8 Flash really cost you.
everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GLM-5.3-Flash cheaper than Gemini 3.8 Flash?
Yes, at current list prices on every line. The example cached session costs $0.16 against $0.71, and 110 sessions a month come to $17.60 against $78.38.
Will Gemini 3.8 Flash get more expensive?
Its current prices are introductory and run through December 31, 2026. Starting January 1, 2027, Google charges $1.50 input, $0.15 cached, and $7.50 output per million tokens for it.
Which model writes longer responses?
GLM-5.3-Flash lists 128K output tokens per response, and Gemini 3.8 Flash 65.5K. Their context windows are 1M and 1.05M.
Is there a free way to try either model?
Gemini 3.8 Flash can be tried on the Gemini API free tier, which includes its input, output, and caching. GLM-5.3-Flash has open weights under the MIT license, so you can run it on your own hardware, at whatever that hardware costs.
Sources
- Z.ai docs: Pricing
- Z.ai docs: GLM-5.3-Flash
- Z.ai: GLM-5.3-Flash
- OpenRouter: GLM-5.3-Flash
- OpenRouter: Rankings
- OpenCode docs: Zen
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Z.ai docs: Context caching
- Google: Context caching