TechCompare LogoTechCompare

GLM-5.3 API Pricing & Cost Calculator

GLM-5.3 keeps GLM-5.2's pricing while improving quality. $4.40/M output with a real 1M context window is the value pick for high-volume document and agentic work against Gemini Flash or GPT-5.6 Terra.

GLM-5.3 is Z.ai's August 2026 flagship refresh, adding better long-context reasoning at the exact same $1.40/$4.40 price point as GLM-5.2, with the 1M context window retained.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

GLM-5.3 is priced at $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's official list, with an ~81% prompt caching discount ($0.26 per million cached reads). There is no published batch discount. That repeats GLM-5.2's rate card exactly and keeps it below Gemini 3.6 Flash's standard $1.50/$7.50 on both rows.

Verdict

The headline $1.40/M input and $4.40/M output with cached reads at $0.26/M matches GLM-5.2 dollar for dollar, so the 5.3 upgrade is a no-cost quality bump for existing GLM workloads. Against Gemini 3.7 Flash's promotional $0.75/$3.75, GLM-5.3 loses on output until the Google promo sunsets at the end of 2026 and lands back at $1.50/$7.50. The 1M context without chunking overhead remains the structural advantage over any model that bills by re-read or retrieval round-trip.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

Is GLM-5.3 priced differently from GLM-5.2?
No. Both are $1.40 per million input, $4.40 per million output, with the same ~81% cache discount ($0.26/M cached reads) and no batch discount. The upgrade from 5.2 to 5.3 is free.
How does GLM-5.3 compare to Gemini 3.7 Flash on price?
Gemini 3.7 Flash is currently on promo at $0.75/$3.75 through December 2026, cheaper than GLM-5.3's $1.40/$4.40. After the promo ends and Flash reverts to $1.50/$7.50, GLM-5.3 wins on both rows. GLM also has the deeper production cache discount story at 81% vs Flash's 90% - Flash still wins on cached input ($0.075/M vs $0.26/M) for stable-prompt loops.
What is the real monthly cost at 50M input and 5M output tokens?
At standard rates with no caching: 50 * $1.40 + 5 * $4.40 = $92. With 50% cache hit rate on input it drops to 50 * $0.83 = $41.50 + $22 = $63.50. Batch discounts aren't published for GLM-5.3, so latency-insensitive workloads should compare against Gemini 3.7 Flash's 50% batch tier.