TechCompare LogoTechCompare

How much VRAM does Nemotron 3.5 Lightning 30B need at Q4_K_M? NVIDIA's single-card 1M model

Nemotron 3.5 Lightning is a 36 GB resident model at the full 1M context, so a 48 GB pro card runs it with all experts loaded. Shorter contexts shrink the cache and the total.

Nemotron 3.5 Lightning 30B at Q4_K_M needs just 36 GB of VRAM with all experts resident at the full 1M context. NVIDIA's hybrid Mamba-2 plus transformer design keeps the cache small, so a 1M-context model fits a single 48 GB card. It only uses 3B active parameters per token, but that doesn't shrink what has to stay loaded, because different tokens pick different experts.

By TechCompare · Updated

Total VRAM required
37.4 GB
Nemotron 3.5 Lightning 30B (MoE) at Q4_K_M
Weights
16.8 GB
30B params
KV cache
17.2 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

37.4 GB

30B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
16.8 GB
KV cache
17.2 GB
Overhead
3.4 GB
Doesn't fit on a 32 GB consumer GPU at Q4_K_M. Q2_K (29.5 GB) is the smallest quant that fits a single RTX 5090.

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

KV cache exceeds model weights: Consider lowering the context length to save on VRAM. Contexts between 8K and 64K are generally more typical for local setups.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

A100 40GB
Datacenter
40 GB
93% used
RTX 6000 Ada
Pro
48 GB
78% used
Apple M3 Max 64GB
Unified
~56 GB
67% of estimate

Just barely too small

RTX 5090
Consumer
32 GB
short by 5.4 GB

How this is calculated

Only 8 attention layers carry a key-value cache. The rest of the transformer blocks are Mamba-2 state-space layers with a fixed-size recurrent state. Weights at Q4_K_M are 16.8 GB, the KV cache with 4 KV heads at head_dim 128 across 8 attention layers is 16.4 GB at 1M, and overhead adds 3.3 GB.

Verdict

The math is what sells NVIDIA's hybrid bet: 16.8 GB of Q4 weights plus a 16.4 GB KV cache (only 8 attention layers contribute, and the Mamba-2 blocks hold a fixed state instead of growing a cache) plus 3.3 GB of overhead totals 36 GB resident. The 3B active count doesn't lower the resident total, because any token can route to any expert. Compared to Nemotron 3 Ultra 550B - the other hybrid on this list - Lightning trades capacity for a footprint a single 48 GB card handles outright.

More Nemotron scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜

Frequently asked questions

How can a 1M-context model fit in 36 GB?
Most of the transformer blocks are Mamba-2 state-space layers, which keep a fixed-size recurrent state instead of growing a key-value cache. Only 8 attention layers accumulate KV. That keeps the cache to 16.4 GB at 1M tokens, next to 16.8 GB of Q4 weights and a few GB of overhead.
Nemotron 3.5 Lightning or gpt-oss 20B for a local agent?
Lightning for long-context agentic work: its 1M native window costs about 36 GB resident. gpt-oss 20B for maximum compatibility and the cleanest tooling, but its window is 128K and it needs about 19.4 GB. Lightning spends much of its budget on cache, gpt-oss on weights.