How much VRAM does Kimi K3 2.8T (MoE) need at Q4_K_M? Moonshot 1M frontier
Kimi K3 2.8T at Q4_K_M is a roughly 1.8 TB resident deployment for 1M context, or about 121 GB if you can tolerate active-expert offload latency.
Kimi K3 2.8T at Q4_K_M with native 1M context needs about 1781 GB of VRAM for an all-resident deployment. Moonshot's third-generation trillion-parameter MoE activates 104B parameters per token from a 2.8 trillion total pool. If you offload cold experts and keep only the hot routed experts in VRAM, the resident footprint drops to roughly 121 GB. The hybrid KDA and Gated-MLA attention keeps the key-value cache surprisingly small at full context.
By TechCompare · Updated
Calculator
Estimated VRAM required
1781 GB
2800B params at Q4_K_M, 1,048,576 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
Hardware that fits
No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.
How this is calculated
The 2.8T parameter pool requires 1568 GB of weights at Q4_K_M regardless of expert routing. Kimi K3's hybrid design uses 69 Kimi Delta Attention layers that store no key-value cache at all, plus 24 Gated-MLA layers that compress per-token KV to a 512-dim latent. Only the 24 MLA layers contribute to the cache, which lands at just 52 GB even at the full 1,048,576 token window. Activation and software overhead adds about 162 GB. In active-only mode the resident weights shrink to 58 GB, which drops total usage to 121 GB. The cache stays 52 GB either way because it depends only on context length, not expert loading.
Verdict
An all-resident Kimi K3 deployment is multi-node datacenter territory: roughly 19 H100 80GB or 11 H200 141GB with NVLink. Active-only fits on dual 80GB cards or a 128GB unified memory workstation, but cold-expert PCIe streaming limits throughput. The KDA plus MLA hybrid is what keeps the cache small enough for active-offload to even be viable at 1M context.
More Kimi scenarios
Frequently asked questions
Why is Kimi K3's key-value cache so small for a 2.8T model?
Can I run Kimi K3 on a single GPU?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Data Read Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜