TechCompare LogoTechCompare

GLM-5.3-Flash API pricing and cost calculator

The direct half-cached 100K-input, 5K-output example costs $0.0115 per request, or $1,150 for 100,000 calls. Cached input is 80% cheaper than uncached input.

GLM-5.3-Flash costs $0.15 per million uncached input tokens and $0.50 per million output tokens on Z.ai's direct API. Cached input costs $0.03 per million tokens. The provider's pricing table lists cache storage as free for a limited time. This page uses those USD rates, checked on October 4, 2026, without importing a reseller promotion or a coding subscription's quota rules. That distinction matters when budgeting a production application that pays for each request.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.

Z.ai: GLM-5.3-Flash
$1.15

How this is calculated

A request with 100,000 input tokens and 5,000 output tokens costs $0.015 for uncached input and $0.0025 for output, totaling $0.0175. Reading half the input from cache reduces input to $0.009 and the request total to $0.0115. At 100 monthly requests, the starting chart is $1.15. At 100,000 identical requests, that is $1,150 with half the input cached or $1,750 without cache hits. A fully cached 100,000-token prompt plus 5,000 generated tokens costs $0.0055. The 80% input discount doesn't apply to output. Output costs become more visible when the model generates long explanations or reasoning traces, so use the provider's billed output usage rather than the length of the final answer alone. Z.ai's model guide lists a 1M-token context window and maximum output of 128K tokens. Those are capacity limits, not a recommended request size. Sending more context raises token volume even when the rate remains the same. These are normal-speed real-time API prices. A coding plan's off-peak points policy describes subscription usage and needs its own calculation. Z.ai lists web search at $0.01 per use, so 100,000 such uses add $1,000 beyond the token estimate.

Verdict

GLM-5.3-Flash is a candidate for coding, document processing, and tasks that use visual input. Test its results on the work you plan to automate and measure complete request usage. Keep reusable instructions near the start of the prompt, check actual cache hits, and give the model a suitable output budget. Compare another provider using the same workload and delivery timing. Cache-heavy requests can rank differently from requests with little reused input, so one uncached example shouldn't decide every deployment.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
DeepSeek V4.1 Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

What does GLM-5.3-Flash cost on Z.ai?
Z.ai lists $0.15 per million uncached input tokens, $0.03 per million cached input tokens, and $0.50 per million output tokens. The chart uses these direct USD rates. Cache storage is listed as free for a limited time, so check the provider's policy before assuming storage will remain free over the life of an application.
How much does prompt caching save?
The cached-input rate is 80% below the uncached rate. With 100,000 input tokens and 5,000 output tokens, reading half the input from cache lowers the total from $0.0175 to $0.0115. At 100,000 requests, that saves $600 on the stated workload. Output remains billed at $0.50 per million tokens regardless of that input cache percentage.
Does a GLM Coding Plan off-peak discount change API rates?
A coding subscription uses its own points and quota rules. The prices on this page are Z.ai's normal-speed real-time pay-as-you-go API rates, so a subscription points discount isn't applied to the chart. Keep account balance spending and plan usage separate when comparing an application budget with the cost of a personal coding subscription.
Which identifier selects GLM-5.3-Flash on Z.ai?
Use glm-5.3-flash as the model code in a request to Z.ai's model API. Check that identifier in your application configuration before applying the rates on this page. The model guide lists a 1M-token context window and maximum output of 128K tokens. Use billed token usage from representative requests when estimating coding, document, or visual workloads.
How should I budget for GLM reasoning and visual tasks?
Use the input and output usage reported for representative requests. Z.ai's model guide says thinking is enabled for GLM-5.3-Flash and can't be disabled, so a short visible answer doesn't guarantee a small output bill. Images and other inputs also need their actual billed usage. Test document and visual workloads separately from a text-only token example.