How much VRAM does Nemotron 3.5 Lightning 30B need at Q4_K_M? NVIDIA's single-card 1M model
Nemotron 3.5 Lightning is a 36 GB resident model at the full 1M context, so a 48 GB pro card runs it with all experts loaded. Shorter contexts shrink the cache and the total.
Nemotron 3.5 Lightning 30B at Q4_K_M needs just 36 GB of VRAM with all experts resident at the full 1M context. NVIDIA's hybrid Mamba-2 plus transformer design keeps the cache small, so a 1M-context model fits a single 48 GB card. It only uses 3B active parameters per token, but that doesn't shrink what has to stay loaded, because different tokens pick different experts.
By TechCompare · Updated
Calculator
Estimated VRAM required
37.4 GB
30B params at Q4_K_M, 1,048,576 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
KV cache exceeds model weights: Consider lowering the context length to save on VRAM. Contexts between 8K and 64K are generally more typical for local setups.
Hardware that fits
Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.
Just barely too small
How this is calculated
Only 8 attention layers carry a key-value cache. The rest of the transformer blocks are Mamba-2 state-space layers with a fixed-size recurrent state. Weights at Q4_K_M are 16.8 GB, the KV cache with 4 KV heads at head_dim 128 across 8 attention layers is 16.4 GB at 1M, and overhead adds 3.3 GB.
Verdict
The math is what sells NVIDIA's hybrid bet: 16.8 GB of Q4 weights plus a 16.4 GB KV cache (only 8 attention layers contribute, and the Mamba-2 blocks hold a fixed state instead of growing a cache) plus 3.3 GB of overhead totals 36 GB resident. The 3B active count doesn't lower the resident total, because any token can route to any expert. Compared to Nemotron 3 Ultra 550B - the other hybrid on this list - Lightning trades capacity for a footprint a single 48 GB card handles outright.
More Nemotron scenarios
Related guides
Frequently asked questions
How can a 1M-context model fit in 36 GB?
Nemotron 3.5 Lightning or gpt-oss 20B for a local agent?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Memory and Storage Throughput Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜