How much VRAM does Nemotron 3.5 Lightning 30B need at Q4_K_M? NVIDIA's single-card 1M model
Nemotron 3.5 Lightning runs the full 1M context on a single 24 GB card with active-only loading (~20 GB), which no comparable 1M-window model manages. Resident mode is 36 GB, so a 48 GB pro card runs it with all experts warm.
Nemotron 3.5 Lightning 30B at Q4_K_M needs just 36 GB of VRAM with all experts resident at the full 1M context, or 20 GB with active-expert offload. NVIDIA's hybrid Mamba-2 plus transformer design with only 3B active parameters per token is what makes a 1M-context model land inside a single 24 GB budget - the reason it's the hottest local agent model on HuggingFace right now.
By TechCompare · Updated
Calculator
Estimated VRAM required
37.4 GB
30B params at Q4_K_M, 1,048,576 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
KV cache exceeds model weights: Consider lowering the context length to save on VRAM. Contexts between 8K and 64K are generally more typical for local setups.
Hardware that fits
Just barely too small
How this is calculated
Only 8 attention layers carry a key-value cache; the rest of the transformer blocks are Mamba-2 state-space layers with a fixed-size recurrent state. Weights at Q4_K_M are 16.8 GB, the KV cache with 4 KV heads at head_dim 128 across 8 attention layers is 16.4 GB at 1M, and overhead adds 3.3 GB. Active-only keeps just the 3B routed experts resident (1.7 GB) so the total drops to ~20 GB, fitting a 24 GB consumer card with room for batch.
Verdict
The math is what sells NVIDIA's hybrid bet: 16.8 GB of Q4 weights plus a 16.4 GB KV cache (only 8 attention layers contribute, and the Mamba-2 blocks hold a fixed state instead of growing a cache) plus 3.3 GB of overhead totals 36 GB resident. Active-only drops weights to 1.7 GB of routed experts and the total to ~20 GB. Compared to Nemotron 3 Ultra 550B - the other hybrid on this list - Lightning trades capacity for a footprint a single RTX 4090 or 5090 handles outright.
More Nemotron scenarios
Related guides
Frequently asked questions
How can a 1M-context model fit in 20 GB?
Nemotron 3.5 Lightning or gpt-oss 20B for a local agent?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Memory and Storage Latency Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜