TechCompare LogoTechCompare

GLM-5.3-Flash VRAM requirements: about 217 GB at Q4 and 1M context

GLM-5.3-Flash needs about 217 GB in this Q4, 1M, batch-one budget. It includes the Transformers sparse-cache layout and fixed KDA state.

GLM-5.3-Flash combines a roughly 320B parameter pool with about 18B parameters active per token. Its attention layout matters as much as its MoE label: 34 layers keep fixed linear-attention state, while 11 sparse layers retain context-dependent storage. This page estimates 179.20 GB of Q4 weights and a total of 216.78 GB using FP16 cache and the Transformers implementation. Other serving engines can allocate the sparse indexer differently.

By TechCompare · Updated

Total VRAM required
217 GB
GLM-5.3-Flash 320B (MoE) at Q4_K_M
Weights
179 GB
320B params
KV cache
17.7 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

217 GB

320B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
179 GB
KV cache
17.7 GB
Overhead
19.8 GB

Estimate accuracy: Planning estimate. Rounded model counts, tensor formats, cache packing, and workspace can change actual allocation. Confirm the exact checkpoint and serving engine with a load test.

GLM-5.3-Flash uses 34 fixed-state KDA layers and 11 sparse MLA layers. This cache budget follows the Transformers implementation, including its unpooled indexer arrays. Optimized serving engines can use smaller caches. Q8 is a sizing scenario and requires backend support.

Hybrid linear attention applied: 34 of 45 layers use fixed recurrent state, so only 11 sparse-attention layers contribute to the context-growing KV cache. The fixed state adds about 0.15 GB at batch 1 and does not grow with context length.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

Apple M3 Ultra 256GB
Unified
~248 GB
87% of estimate
NVIDIA B300
Datacenter
288 GB
75% used

Just barely too small

MI300X
Datacenter
192 GB
short by 24.8 GB

How this is calculated

The configuration specifies 45 layers, hidden size 4,096, 64 KDA heads of width 128, and a four-position convolution kernel. Each KDA layer retains a 64 x 128 x 128 FP32 recurrent state plus a BF16 convolution state for three projections. Across 34 layers that is about 0.149 GB at batch one, independent of context length. The sparse layers cache a shared 512-value MLA latent rather than separate expanded query-head keys and values. The linked Transformers code also stores unpooled indexer keys, gate scores, and validity values, adding 128 + 128 + 1 values per token in each sparse layer. Its total logical cache budget is therefore 11 x 769 values per input token, or 17.74 GB in FP16 at 1,048,576 tokens. Including 10% general overhead and fixed state gives 216.78 GB, about 201.89 GiB. At 32K context the total is 197.88 GB. An optimized pooled indexer can reduce this cache allocation.

Verdict

Use this as an implementation-specific planning budget, then compare it with your serving engine's measured allocation. Linear attention keeps most layers from accumulating a full sequence cache, but it doesn't make the 320B weight pool disappear. The Q4 scenario needs far more memory than one conventional 80 GB accelerator even at 32K. Keep room for vision activations, prompt-processing workspace, and per-shard imbalance when evaluating a multi-GPU setup. Q4 or Q8 cache availability must be checked for that exact backend.

More GLM scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
DeepSeek V4.1 Flash Q4 + FP8 Engram
Include 552B backbone weights and the separate 196B Engram table.
View details ➜

Frequently asked questions

Why doesn't the calculator use all 64 attention heads for KV?
The sparse path caches the compressed MLA latent and expands it for attention computation. Counting a separate full-width cache for every query head would overstate the retained storage. This preset uses one shared latent per token per sparse layer and separately includes indexer arrays. Query projection dimensions describe work performed by attention, not necessarily tensors that remain in the cache.
Does sparse attention store only the selected tokens?
Not in the implementation modeled here. Selecting a small set of positions reduces attention work, but the cache retains latent states for the sequence and the indexer keeps information used to choose those positions. The preset follows the linked Transformers cache writes. A runtime that pools indexer storage or uses a different packed layout can have a smaller cache, so a benchmark should identify its engine and revision.
How does batch size affect the linear-attention state?
Each concurrent sequence needs its own recurrent and convolution state. Doubling batch doubles that fixed state and the sequence cache, while resident weights are shared. The fixed state doesn't double when you double one sequence's context length. This distinction is why treating all 45 layers as ordinary full attention would give an incorrect memory curve for GLM Flash.
Are 320B and 18B exact checkpoint byte counts?
They're the publisher's rounded total and active parameter figures. The preset uses 320B as a weight planning input and doesn't infer tensor bytes from the 18B active count. Actual allocation depends on quantized tensor blocks, scales, modules that retain higher precision, and optional components. Inspect the checkpoint's tensor inventory and loaded memory to reconcile a deployment with this rounded estimate.
Will changing reasoning effort change resident model size?
Reasoning effort changes the generation workload rather than this architecture's stored weight pool. Longer outputs can grow context and occupy a request slot for longer, increasing cache usage and concurrency pressure. Size the service around combined prompt and output length, not only average input size. The publisher documents low, high, and max effort settings, so load tests should use the effort level you'll actually serve.