How much VRAM does Muse Glimmer 30B need at Q4_K_M? Native 131K context
Muse Glimmer 30B at Q4_K_M fits a 24 GB card at its native 131K context at about 20.5 GB. The calculator also puts an extended 256K run near 22.4 GB, but 256K should not be described as native until the serving stack and model release document that support.
Muse Glimmer 30B at Q4_K_M needs about 20.5 GB of VRAM at its native 131K context: 16.8 GB of weights, 1.8 GB of KV cache, and 1.9 GB of overhead. The official architecture has 52 layers with three local attention layers for every global layer, so the 256K calculator scenario is an extended run rather than native support. Apache 2.0 licensing and a sub-24 GB native estimate make it a practical single-GPU model.
By TechCompare · Updated
Calculator
Estimated VRAM required
20.0 GB
30B params at Q4_K_M, 131,072 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
Sliding-window attention applied: This model caps 4 of every 5 layers at a 1024-token window. KV cache estimate is 80% smaller than naive full-attention math at this context length.
Hardware that fits
Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.
How this is calculated
The released config uses 52 layers with a [Local, Local, Local, Global] pattern. Three out of every four layers cap their KV cache at a 2048-token sliding window, while 13 global layers retain the full sequence. With 2 KV heads and head_dim 128, that gives about 1.83 GB of FP16 KV cache at native 131,072 tokens and a calculated total of 20.5 GB. At an extended 262,144-token scenario, the hybrid cache is about 3.57 GB and the total is about 22.4 GB. Using full-context KV storage for all 52 layers would be about 3.49 GB at native 131K, not the old ~100 GB comparison. The 30B total at Q4_K_M is 16.8 GB of weights, and Apache 2.0 licensing means no commercial strings.
Verdict
The corrected native estimate is 20.5 GB: 16.8 GB of Q4 weights, 1.83 GB of FP16 KV cache from 13 global layers plus 39 local layers, and 1.86 GB of overhead. At an extended 256K scenario, the same formula gives 22.4 GB with a 3.57 GB cache. That moves the RTX 3090 into the fit list for native context and the modeled extended scenario, while keeping the important limit clear: 131K is the released native position limit in the official config. The comparison that matters is Gemma 4 31B dense at a similar footprint but with a Google license, versus Glimmer with Apache 2.0.
More Meta scenarios
Related guides
Frequently asked questions
Is Muse Glimmer the successor to Llama 4?
Muse Glimmer 30B or Qwen 3.8 27B?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Memory and Storage Throughput Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜