Model comparison
GLM-5.3 vs Gemini 3.1 Pro Preview: output price decides it
Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.
· Prices as of September 28, 2026
GLM-5.3
Z.ai · Released August 14, 2026
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
GLM-5.3 facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
GLM-5.3 costs $1.44 on the example agentic coding session against $2.00 on Gemini 3.1 Pro Preview, 28% less, and the gap grows to 2.6x on output-heavy work because Gemini output costs $12 per million against $4.40. Pick Gemini 3.1 Pro Preview if you work in Gemini CLI, where it is the Pro half of the default auto model; pick GLM-5.3 for cheaper output, a 128K output limit, and open weights.
Choose GLM-5.3 if
- Your work is output-heavy: the example generation costs $0.39 on GLM-5.3 against $1.02 on Gemini 3.1 Pro Preview.
- You need long single responses, since GLM-5.3 writes up to 128K tokens per request against 65.5K.
- You want a current release with open weights rather than a preview, under Z.ai's own GLM-5.3 license.
Choose Gemini 3.1 Pro Preview if
- You work in Gemini CLI, where Gemini 3.1 Pro Preview is the Pro half of the default auto model.
- Your sessions reread a stable prefix, where a Gemini cache hit costs $0.20 per million against $0.26 on GLM-5.3, though its writes, input, and output still cost more.
- You want Google's customtools endpoint for Gemini 3.1 Pro Preview, which Google says is better at prioritizing custom tools alongside bash, at the same price.
- You pick models in Cursor, which offers Gemini 3.1 Pro Preview and does not offer GLM-5.3.
Side by side
Specs and prices
| Fact | GLM-5.3 | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | Z.ai | |
| API model id | glm-5.3 | gemini-3.1-pro-preview |
| Released | August 14, 2026 | February 19, 2026 |
| Status | Current | Preview |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $1.40 | $2 |
| Cache hit, per 1M | $0.26 | $0.20 |
| Cache write, per 1M | $1.40 (same as input) | $2 (same as input) |
| Output, per 1M tokens | $4.40 | $12 |
| Runs in | OpenCode and OpenRouter | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GLM-5.3 | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.44 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.25 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $0.39 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $158.40 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.56 | $0.80 |
| Cache reads | $0.52 | $0.40 |
| Uncached input | $0.14 | $0.20 |
| Output | $0.22 | $0.60 |
- caching saves on the session with GLM-5.3 (61%)
- $2.28
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
Neither maker charges a cache-write premium
Z.ai and Google bill cache writes the same way: written tokens cost ordinary input. GLM-5.3 writes at $1.40 per million and Gemini 3.1 Pro Preview at $2. That is the same 1.4x ratio as their input prices. In the example session, the 400K written tokens cost $0.56 and $0.80.
Cache hits favor Gemini. A hit costs 10% of the input price on Gemini 3.1 Pro Preview, $0.20 per million, and $0.26 on GLM-5.3, which is 18.6% of its input. The session's 2M cached tokens cost $0.40 on Gemini and $0.52 on GLM-5.3, so the cheaper model spends $0.12 more on reads.
Output decides the result. Gemini charges $12 per million output tokens and GLM-5.3 charges $4.40, a 2.7x spread. The session writes only 50K tokens of output, yet that line is $0.60 on Gemini, 30% of its total, against $0.22 on GLM-5.3. Output accounts for $0.38 of the $0.56 difference, and the session ends at $2.00 against $1.44.
The gap on output-heavy work and over a month
When output dominates, the ratio widens. The output-heavy generation, 30K tokens in and 80K out, costs $1.02 on Gemini 3.1 Pro Preview and $0.39 on GLM-5.3. The large one-off review, which is mostly input, costs $0.42 against $0.25.
Across 110 sessions a month the session comes to $220.00 on Gemini and $158.40 on GLM-5.3, $61.60 apart. Caching saves 64% on Gemini and 61% on GLM-5.3 against sending the same tokens uncached. These are API-equivalent estimates at each maker's own rates, and OpenRouter providers serving GLM-5.3's open weights can charge differently from Z.ai.
Thinking adds output. Gemini 3.1 Pro Preview's thinking cannot be turned off and defaults to high. GLM-5.3 always reasons too, at low, high, or max. GLM-5.3 and Gemini also use different tokenizers, so the same code will not count as the same number of tokens on both.
Output limits, long prompts, and preview status
GLM-5.3 writes up to 128K tokens in one response, about 2x the 65.5K limit on Gemini 3.1 Pro Preview. That gap matters for agents that emit whole files or long plans in a single turn. Context is close: 1M on GLM-5.3 and 1.05M on Gemini.
Gemini 3.1 Pro Preview raises its rates on prompts over 200K input tokens, to $4 input, $0.40 cached, and $18 output per million. The GLM-5.3 price in this comparison lists no such tier. The example session stays under 200K per request, so Gemini's higher tier never applies to it.
Gemini 3.1 Pro Preview is, as its name says, a preview. Google has announced Gemini 3.5 Pro, which is not yet released, and GitHub Copilot retired Gemini 3.1 Pro Preview on September 1, 2026. GLM-5.3 is a current release from August 14, 2026.
Where each runs and how Google and Z.ai pitch them
Gemini CLI is Google's own coding agent. Gemini 3.1 Pro Preview is the Pro half of its default auto model, and Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint at the same price. Cursor, OpenCode, and OpenRouter offer the model too. Google describes it as having "advanced intelligence, complex problem-solving skills, and powerful agentic and vibe coding capabilities."
GLM-5.3 is offered through OpenRouter and in OpenCode, and Z.ai's API accepts OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages requests. Z.ai calls it its latest flagship, "delivering comprehensive advancements in complex software engineering and agent capabilities."
Google's caching carries one more decision. On Gemini 3.1 Pro Preview, implicit caching discounts a repeated prefix on its own, though Google does not promise a hit. Its explicit caching trades that uncertainty for a named cache with a fixed discount, plus storage at $4.50 per million tokens per hour on Pro models. GLM-5.3's caching is automatic, and Z.ai says cached input storage is free for a limited time.
Prompt caching
How each maker bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what GLM-5.3 and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GLM-5.3 cheaper than Gemini 3.1 Pro Preview?
Yes, on every example workload: $1.44 against $2.00 for the agentic session, $0.25 against $0.42 for the uncached review, and $0.39 against $1.02 for output-heavy generation. Gemini charges less for a cache hit, $0.20 against $0.26 per million.
Does either model charge extra to write the cache?
No. Z.ai lists no write fee, and Google publishes no separate write price, so written tokens cost ordinary input on both. Google's explicit caching on Gemini 3.1 Pro Preview does add a storage charge for as long as the cache lives.
Which model can write longer responses?
GLM-5.3, with a 128K output limit against 65.5K on Gemini 3.1 Pro Preview. The context windows are close, at 1M and 1.05M tokens.
How can I compare what each model costs me?
EveryToken prices Gemini 3.1 Pro Preview at Google's rates from Gemini CLI, Cursor, and OpenCode history, and prices GLM-5.3 from OpenRouter's catalog when you use it through OpenRouter. It does not price GLM-5.3 called directly on Z.ai's API.
Sources
- Z.ai docs: Pricing
- Z.ai docs: GLM-5.3
- Z.ai: GLM-5.3
- OpenRouter: GLM-5.3
- OpenCode docs: Zen
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Z.ai docs: Context caching
- Google: Context caching