How much VRAM does Nemotron 3 Ultra 550B A55B (MoE) need at Q4_K_M? NVIDIA 1M Mamba-2
Nemotron 3 Ultra 550B at Q4_K_M is a 353 GB resident deployment for 1M context, or about 48 GB with active-expert offload. The Mamba-2 layers are why the cache is tiny enough for active-only to almost fit on a single card.
Nemotron 3 Ultra 550B at Q4_K_M with native 1M context needs about 353 GB of VRAM with all experts resident. NVIDIA's hybrid Mamba-2 plus MoE flagship activates 55B parameters per token from a 550B total pool. Only 12 of 108 transformer blocks carry attention and the other 48 are Mamba-2 state-space layers with zero key-value cache. Active-expert offload drops the resident footprint to roughly 48 GB.
By TechCompare · Updated
Calculator
Estimated VRAM required
353 GB
550B params at Q4_K_M, 1,048,576 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
Hardware that fits
No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.
How this is calculated
The 550B parameter pool requires 308 GB of weights at Q4_K_M. The architecture is a hybrid: 12 attention layers, 48 Mamba-2 state-space layers, and 48 MoE expert blocks. Only the 12 attention layers contribute to the key-value cache, which is just 13 GB at the full 1M context with 2 KV heads at head_dim 128. The Mamba-2 layers maintain a fixed recurrent state instead of growing a cache. Activation and software overhead adds about 32 GB. In active-only mode the resident weights shrink to 31 GB, which drops total usage to 48 GB. The tiny cache means active-only is remarkably efficient for a 550B model.
Verdict
Resident mode needs five 80GB datacenter cards or three 141GB H200 cards. Active-only at 48 GB fits on a single 80GB card with headroom for batch. The 48 Mamba-2 layers replace what would otherwise be a 100 GB-plus key-value cache at 1M context, which is why this is the most single-card-friendly 500B-class MoE on the list.
More Nemotron scenarios
Frequently asked questions
Why is Nemotron 3 Ultra's key-value cache so small?
How does Nemotron 3 Ultra compare to Nemotron 3 Super?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Data Read Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜