GLM-5.3-Flash VRAM requirements: about 217 GB at Q4 and 1M context
GLM-5.3-Flash needs about 217 GB in this Q4, 1M, batch-one budget. It includes the Transformers sparse-cache layout and fixed KDA state.
GLM-5.3-Flash combines a roughly 320B parameter pool with about 18B parameters active per token. Its attention layout matters as much as its MoE label: 34 layers keep fixed linear-attention state, while 11 sparse layers retain context-dependent storage. This page estimates 179.20 GB of Q4 weights and a total of 216.78 GB using FP16 cache and the Transformers implementation. Other serving engines can allocate the sparse indexer differently.
By TechCompare · Updated
Calculator
Estimated VRAM required
217 GB
320B params at Q4_K_M, 1,048,576 token context, batch 1, inference.
Estimate accuracy: Planning estimate. Rounded model counts, tensor formats, cache packing, and workspace can change actual allocation. Confirm the exact checkpoint and serving engine with a load test.
GLM-5.3-Flash uses 34 fixed-state KDA layers and 11 sparse MLA layers. This cache budget follows the Transformers implementation, including its unpooled indexer arrays. Optimized serving engines can use smaller caches. Q8 is a sizing scenario and requires backend support.
Hybrid linear attention applied: 34 of 45 layers use fixed recurrent state, so only 11 sparse-attention layers contribute to the context-growing KV cache. The fixed state adds about 0.15 GB at batch 1 and does not grow with context length.
Hardware that fits
Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.
Just barely too small
How this is calculated
The configuration specifies 45 layers, hidden size 4,096, 64 KDA heads of width 128, and a four-position convolution kernel. Each KDA layer retains a 64 x 128 x 128 FP32 recurrent state plus a BF16 convolution state for three projections. Across 34 layers that is about 0.149 GB at batch one, independent of context length. The sparse layers cache a shared 512-value MLA latent rather than separate expanded query-head keys and values. The linked Transformers code also stores unpooled indexer keys, gate scores, and validity values, adding 128 + 128 + 1 values per token in each sparse layer. Its total logical cache budget is therefore 11 x 769 values per input token, or 17.74 GB in FP16 at 1,048,576 tokens. Including 10% general overhead and fixed state gives 216.78 GB, about 201.89 GiB. At 32K context the total is 197.88 GB. An optimized pooled indexer can reduce this cache allocation.
Verdict
Use this as an implementation-specific planning budget, then compare it with your serving engine's measured allocation. Linear attention keeps most layers from accumulating a full sequence cache, but it doesn't make the 320B weight pool disappear. The Q4 scenario needs far more memory than one conventional 80 GB accelerator even at 32K. Keep room for vision activations, prompt-processing workspace, and per-shard imbalance when evaluating a multi-GPU setup. Q4 or Q8 cache availability must be checked for that exact backend.
More GLM scenarios
Related guides
Frequently asked questions
Why doesn't the calculator use all 64 attention heads for KV?
Does sparse attention store only the selected tokens?
How does batch size affect the linear-attention state?
Are 320B and 18B exact checkpoint byte counts?
Will changing reasoning effort change resident model size?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Memory and Storage Throughput Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜