TechCompare LogoTechCompare

How much VRAM does Muse Glimmer 30B need at Q4_K_M? Native 131K context

Muse Glimmer 30B at Q4_K_M fits a 24 GB card at its native 131K context at about 20.5 GB. The calculator also puts an extended 256K run near 22.4 GB, but 256K should not be described as native until the serving stack and model release document that support.

Muse Glimmer 30B at Q4_K_M needs about 20.5 GB of VRAM at its native 131K context: 16.8 GB of weights, 1.8 GB of KV cache, and 1.9 GB of overhead. The official architecture has 52 layers with three local attention layers for every global layer, so the 256K calculator scenario is an extended run rather than native support. Apache 2.0 licensing and a sub-24 GB native estimate make it a practical single-GPU model.

By TechCompare · Updated

Total VRAM required
20.5 GB
Muse Glimmer 30B at Q4_K_M
Weights
16.8 GB
30B params
KV cache
1.8 GB
128K tokens, FP16 KV

Calculator

Estimated VRAM required

20.0 GB

30B params at Q4_K_M, 131,072 token context, batch 1, inference.

Weights
16.8 GB
KV cache
1.4 GB
Overhead
1.8 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Sliding-window attention applied: This model caps 4 of every 5 layers at a 1024-token window. KV cache estimate is 80% smaller than naive full-attention math at this context length.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

RTX 3090
Consumer
24 GB
83% used
RTX 5090
Consumer
32 GB
63% used
A100 40GB
Datacenter
40 GB
50% used
Apple M3 Max 64GB
Unified
~56 GB
36% of estimate

How this is calculated

The released config uses 52 layers with a [Local, Local, Local, Global] pattern. Three out of every four layers cap their KV cache at a 2048-token sliding window, while 13 global layers retain the full sequence. With 2 KV heads and head_dim 128, that gives about 1.83 GB of FP16 KV cache at native 131,072 tokens and a calculated total of 20.5 GB. At an extended 262,144-token scenario, the hybrid cache is about 3.57 GB and the total is about 22.4 GB. Using full-context KV storage for all 52 layers would be about 3.49 GB at native 131K, not the old ~100 GB comparison. The 30B total at Q4_K_M is 16.8 GB of weights, and Apache 2.0 licensing means no commercial strings.

Verdict

The corrected native estimate is 20.5 GB: 16.8 GB of Q4 weights, 1.83 GB of FP16 KV cache from 13 global layers plus 39 local layers, and 1.86 GB of overhead. At an extended 256K scenario, the same formula gives 22.4 GB with a 3.57 GB cache. That moves the RTX 3090 into the fit list for native context and the modeled extended scenario, while keeping the important limit clear: 131K is the released native position limit in the official config. The comparison that matters is Gemma 4 31B dense at a similar footprint but with a Google license, versus Glimmer with Apache 2.0.

More Meta scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜

Frequently asked questions

Is Muse Glimmer the successor to Llama 4?
It's Meta's 2026 open release rather than a numbered Llama sequel. Glimmer is multimodal (28B text + 2B vision), Apache 2.0, and built around hybrid sliding-window attention. There's no confirmed Llama 5 open-weight release as of this writing.
Muse Glimmer 30B or Qwen 3.8 27B?
Glimmer if you want Apache 2.0 licensing and a smaller native footprint (about 20.5 GB at 131K, or about 22.4 GB in the calculator's extended 256K scenario). Qwen 3.8 27B if you need its longer native window and can budget the much larger long-context footprint. Pick on the context limit, license, and ecosystem you actually need.