GLM-5.3-Flash API pricing and cost calculator
The direct half-cached 100K-input, 5K-output example costs $0.0115 per request, or $1,150 for 100,000 calls. Cached input is 80% cheaper than uncached input.
GLM-5.3-Flash costs $0.15 per million uncached input tokens and $0.50 per million output tokens on Z.ai's direct API. Cached input costs $0.03 per million tokens. The provider's pricing table lists cache storage as free for a limited time. This page uses those USD rates, checked on October 4, 2026, without importing a reseller promotion or a coding subscription's quota rules. That distinction matters when budgeting a production application that pays for each request.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.
How this is calculated
A request with 100,000 input tokens and 5,000 output tokens costs $0.015 for uncached input and $0.0025 for output, totaling $0.0175. Reading half the input from cache reduces input to $0.009 and the request total to $0.0115. At 100 monthly requests, the starting chart is $1.15. At 100,000 identical requests, that is $1,150 with half the input cached or $1,750 without cache hits. A fully cached 100,000-token prompt plus 5,000 generated tokens costs $0.0055. The 80% input discount doesn't apply to output. Output costs become more visible when the model generates long explanations or reasoning traces, so use the provider's billed output usage rather than the length of the final answer alone. Z.ai's model guide lists a 1M-token context window and maximum output of 128K tokens. Those are capacity limits, not a recommended request size. Sending more context raises token volume even when the rate remains the same. These are normal-speed real-time API prices. A coding plan's off-peak points policy describes subscription usage and needs its own calculation. Z.ai lists web search at $0.01 per use, so 100,000 such uses add $1,000 beyond the token estimate.
Verdict
GLM-5.3-Flash is a candidate for coding, document processing, and tasks that use visual input. Test its results on the work you plan to automate and measure complete request usage. Keep reusable instructions near the start of the prompt, check actual cache hits, and give the model a suitable output budget. Compare another provider using the same workload and delivery timing. Cache-heavy requests can rank differently from requests with little reused input, so one uncached example shouldn't decide every deployment.
More API Standalones scenarios
Related guides
Frequently asked questions
What does GLM-5.3-Flash cost on Z.ai?
How much does prompt caching save?
Does a GLM Coding Plan off-peak discount change API rates?
Which identifier selects GLM-5.3-Flash on Z.ai?
How should I budget for GLM reasoning and visual tasks?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with readable output and error detection.
Use tool ➜