How much VRAM does Muse Glimmer 30B need at Q4_K_M? Meta's open-weights return
Muse Glimmer 30B at Q4_K_M fits one 32 GB card at its full 256K context (28 GB total), and one 24 GB card if you cap context at 128K. It's the pick when you want Meta-quality weights with Llama-clean licensing in a single-GPU footprint.
Muse Glimmer 30B at Q4_K_M needs about 28 GB of VRAM at its native 256K context - 17 GB of weights plus an 8.7 GB cache kept small by hybrid sliding-window attention plus 2.6 GB overhead. It's Meta's first major open release since the Llama 4 line, Apache 2.0 licensed, and sized for a single GPU.
By TechCompare · Updated
Calculator
Estimated VRAM required
28.1 GB
30B params at Q4_K_M, 262,144 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
Sliding-window attention applied: This model caps 4 of every 5 layers at a 1024-token window. KV cache estimate is 80% smaller than naive full-attention math at this context length.
Hardware that fits
Just barely too small
How this is calculated
Glimmer alternates full-attention and 1024-token sliding-window layers, with full attention every 5th layer. That's the Gemma-style recipe that keeps the KV cache at 8.7 GB at 256K instead of the ~100 GB a plain full-attention 40-layer model would burn at the same window. The 30B total (28B text + 2B vision) at Q4_K_M is 16.8 GB of weights, and Apache 2.0 licensing means no commercial strings.
Verdict
The 28 GB resident total is 16.8 GB of Q4 weights, 8.7 GB of FP16 KV cache from the hybrid SWA recipe (full attention every 5th layer, 1024-token window elsewhere), and 2.6 GB of overhead. A 32 GB card or a 4090 with 128K context cap runs it cleanly. The comparison that matters: Gemma 4 31B dense at the same class footprint but with a Google license, versus Glimmer with Apache 2.0, and Qwen 3.8 27B with a longer native window but a cache that punishes long contexts.
More Meta scenarios
Related guides
Frequently asked questions
Is Muse Glimmer the successor to Llama 4?
Muse Glimmer 30B or Qwen 3.8 27B?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Memory and Storage Latency Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜