TechCompare LogoTechCompare

GLM-5.2 API Pricing & Cost Calculator

GLM-5.2 is the value workhorse of the Chinese frontier labs. Excellent output pricing ($4.40 per million) and a real 1M context window make it a strong pick for high-volume agentic and document workloads.

GLM-5.2 is Z.ai's (Zhipu) frontier model, offering a 1-million-token context window at prices competitive with Gemini 3.6 Flash and below most US frontier peers.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

GLM-5.2 is priced at $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's official list, with an ~81% prompt caching discount ($0.26 per million cached reads). That puts it cheaper than Gemini 3.6 Flash ($1.50/$7.50) on both axes and below GPT-5.6 Terra's list price on input.

Verdict

GLM-5.2's $1.40/M input and $4.40/M output sit below Gemini 3.6 Flash ($1.50/$7.50) on both rows, with the ~81% cache discount dropping cached input to $0.26/M. The 1M context window is what makes the value case: long-document analysis, agentic loops with large persistent context blocks, and RAG over substantial corpora fit in a single call without chunking. The trade against Gemini Flash is the shallower cache discount (81% vs 90%) and no published batch discount, so workloads with very high cache hit rates or batch-mode jobs may still favor Flash, but for output-heavy text and document work GLM-5.2's $4.40 output row wins on absolute dollar cost.

More API Standalones scenarios

GPT-5.5 Pricing
Single-model gpt-5.5 cost estimate
View details ➜
GPT-5.4 Pricing
Single-model gpt-5.4 cost estimate
View details ➜
Claude Opus 4.8 Pricing
Single-model claude-opus-4.8 cost estimate
View details ➜

Frequently asked questions

How does GLM-5.2 compare to Gemini 3.6 Flash on price?
GLM-5.2 is cheaper on both input ($1.40 vs $1.50 per million) and output ($4.40 vs $7.50 per million). Gemini 3.6 Flash has a stronger cache discount (90% vs ~81%), so the gap narrows for cache-heavy workloads.
How much does GLM-5.2 cost per million output tokens?
GLM-5.2 charges $4.40 per million output tokens, beating Gemini 3.6 Flash's $7.50/M by about 41% on output cost. Its 81% prompt caching discount brings cached input to $0.26 per million, less aggressive than Gemini Flash's 90% ($0.15/M). There is no batch discount on GLM-5.2.