Skip to content

Model comparison

GLM-5.3 vs Gemini 3.1 Pro Preview: output price decides it

Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.

· Prices as of September 28, 2026

  • GLM-5.3

    Z.ai · Released August 14, 2026

    Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

    GLM-5.3 facts and comparisons
  • Gemini 3.1 Pro Preview

    Google · Released February 19, 2026 · Preview

    Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.

    Gemini 3.1 Pro Preview facts and comparisons

The short answer

GLM-5.3 costs $1.44 on the example agentic coding session against $2.00 on Gemini 3.1 Pro Preview, 28% less, and the gap grows to 2.6x on output-heavy work because Gemini output costs $12 per million against $4.40. Pick Gemini 3.1 Pro Preview if you work in Gemini CLI, where it is the Pro half of the default auto model; pick GLM-5.3 for cheaper output, a 128K output limit, and open weights.

Choose GLM-5.3 if

  • Your work is output-heavy: the example generation costs $0.39 on GLM-5.3 against $1.02 on Gemini 3.1 Pro Preview.
  • You need long single responses, since GLM-5.3 writes up to 128K tokens per request against 65.5K.
  • You want a current release with open weights rather than a preview, under Z.ai's own GLM-5.3 license.

Choose Gemini 3.1 Pro Preview if

  • You work in Gemini CLI, where Gemini 3.1 Pro Preview is the Pro half of the default auto model.
  • Your sessions reread a stable prefix, where a Gemini cache hit costs $0.20 per million against $0.26 on GLM-5.3, though its writes, input, and output still cost more.
  • You want Google's customtools endpoint for Gemini 3.1 Pro Preview, which Google says is better at prioritizing custom tools alongside bash, at the same price.
  • You pick models in Cursor, which offers Gemini 3.1 Pro Preview and does not offer GLM-5.3.

Side by side

Specs and prices

FactGLM-5.3Gemini 3.1 Pro Preview
MakerZ.aiGoogle
API model idglm-5.3gemini-3.1-pro-preview
ReleasedAugust 14, 2026February 19, 2026
StatusCurrentPreview
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsYesNo
Input, per 1M tokens$1.40$2
Cache hit, per 1M$0.26$0.20
Cache write, per 1M$1.40 (same as input)$2 (same as input)
Output, per 1M tokens$4.40$12
Runs inOpenCode and OpenRouterCursor, Gemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3Gemini 3.1 Pro Preview
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.44$2.00
Large one-off review, 150K input with no cache hits, 10K output$0.25$0.42
Output-heavy generation, 30K input, 80K output$0.39$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$158.40$220.00
Where the session’s cost goes
Cache writes$0.56$0.80
Cache reads$0.52$0.40
Uncached input$0.14$0.20
Output$0.22$0.60
caching saves on the session with GLM-5.3 (61%)
$2.28
caching saves on the session with Gemini 3.1 Pro Preview (64%)
$3.60

Neither maker charges a cache-write premium

Z.ai and Google bill cache writes the same way: written tokens cost ordinary input. GLM-5.3 writes at $1.40 per million and Gemini 3.1 Pro Preview at $2. That is the same 1.4x ratio as their input prices. In the example session, the 400K written tokens cost $0.56 and $0.80.

Cache hits favor Gemini. A hit costs 10% of the input price on Gemini 3.1 Pro Preview, $0.20 per million, and $0.26 on GLM-5.3, which is 18.6% of its input. The session's 2M cached tokens cost $0.40 on Gemini and $0.52 on GLM-5.3, so the cheaper model spends $0.12 more on reads.

Output decides the result. Gemini charges $12 per million output tokens and GLM-5.3 charges $4.40, a 2.7x spread. The session writes only 50K tokens of output, yet that line is $0.60 on Gemini, 30% of its total, against $0.22 on GLM-5.3. Output accounts for $0.38 of the $0.56 difference, and the session ends at $2.00 against $1.44.

The gap on output-heavy work and over a month

When output dominates, the ratio widens. The output-heavy generation, 30K tokens in and 80K out, costs $1.02 on Gemini 3.1 Pro Preview and $0.39 on GLM-5.3. The large one-off review, which is mostly input, costs $0.42 against $0.25.

Across 110 sessions a month the session comes to $220.00 on Gemini and $158.40 on GLM-5.3, $61.60 apart. Caching saves 64% on Gemini and 61% on GLM-5.3 against sending the same tokens uncached. These are API-equivalent estimates at each maker's own rates, and OpenRouter providers serving GLM-5.3's open weights can charge differently from Z.ai.

Thinking adds output. Gemini 3.1 Pro Preview's thinking cannot be turned off and defaults to high. GLM-5.3 always reasons too, at low, high, or max. GLM-5.3 and Gemini also use different tokenizers, so the same code will not count as the same number of tokens on both.

Output limits, long prompts, and preview status

GLM-5.3 writes up to 128K tokens in one response, about 2x the 65.5K limit on Gemini 3.1 Pro Preview. That gap matters for agents that emit whole files or long plans in a single turn. Context is close: 1M on GLM-5.3 and 1.05M on Gemini.

Gemini 3.1 Pro Preview raises its rates on prompts over 200K input tokens, to $4 input, $0.40 cached, and $18 output per million. The GLM-5.3 price in this comparison lists no such tier. The example session stays under 200K per request, so Gemini's higher tier never applies to it.

Gemini 3.1 Pro Preview is, as its name says, a preview. Google has announced Gemini 3.5 Pro, which is not yet released, and GitHub Copilot retired Gemini 3.1 Pro Preview on September 1, 2026. GLM-5.3 is a current release from August 14, 2026.

Where each runs and how Google and Z.ai pitch them

Gemini CLI is Google's own coding agent. Gemini 3.1 Pro Preview is the Pro half of its default auto model, and Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint at the same price. Cursor, OpenCode, and OpenRouter offer the model too. Google describes it as having "advanced intelligence, complex problem-solving skills, and powerful agentic and vibe coding capabilities."

GLM-5.3 is offered through OpenRouter and in OpenCode, and Z.ai's API accepts OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages requests. Z.ai calls it its latest flagship, "delivering comprehensive advancements in complex software engineering and agent capabilities."

Google's caching carries one more decision. On Gemini 3.1 Pro Preview, implicit caching discounts a repeated prefix on its own, though Google does not promise a hit. Its explicit caching trades that uncertainty for a named cache with a fixed discount, plus storage at $4.50 per million tokens per hour on Pro models. GLM-5.3's caching is automatic, and Z.ai says cached input storage is free for a limited time.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GLM-5.3 and Gemini 3.1 Pro Preview really cost you.

everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3 cheaper than Gemini 3.1 Pro Preview?

Yes, on every example workload: $1.44 against $2.00 for the agentic session, $0.25 against $0.42 for the uncached review, and $0.39 against $1.02 for output-heavy generation. Gemini charges less for a cache hit, $0.20 against $0.26 per million.

Does either model charge extra to write the cache?

No. Z.ai lists no write fee, and Google publishes no separate write price, so written tokens cost ordinary input on both. Google's explicit caching on Gemini 3.1 Pro Preview does add a storage charge for as long as the cache lives.

Which model can write longer responses?

GLM-5.3, with a 128K output limit against 65.5K on Gemini 3.1 Pro Preview. The context windows are close, at 1M and 1.05M tokens.

How can I compare what each model costs me?

EveryToken prices Gemini 3.1 Pro Preview at Google's rates from Gemini CLI, Cursor, and OpenCode history, and prices GLM-5.3 from OpenRouter's catalog when you use it through OpenRouter. It does not price GLM-5.3 called directly on Z.ai's API.

  • Gemini 3.1 Pro Preview vs Gemini 2.5 Pro

    Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • Claude Fable 5.1 vs Gemini 3.1 Pro Preview

    A cached coding session costs 5.3x more on Claude Fable 5.1 than on Gemini 3.1 Pro Preview, mostly from cache writes. Output limits and preview status compared.

  • Claude Opus 4.8 vs Gemini 3.1 Pro Preview

    Claude Opus 4.8 costs 3x Gemini 3.1 Pro Preview on a cached coding session, $6.00 against $2.00. How long prompts, fast mode, and preview status shift that.

  • Claude Opus 5.5 vs Gemini 3.1 Pro Preview

    Gemini 3.1 Pro Preview costs less than half of Claude Opus 5.5 on a cached coding session, though both charge $0.20 per cache hit. Where the gap comes from.