How much VRAM does Kimi K3 2.8T (MoE) need at Q4_K_M? Moonshot 1M frontier

Kimi K3 2.8T at Q4_K_M is a roughly 1.8 TB resident deployment for 1M context, or about 121 GB if you can tolerate active-expert offload latency.

Kimi K3 2.8T at Q4_K_M with native 1M context needs about 1781 GB of VRAM for an all-resident deployment. Moonshot's third-generation trillion-parameter MoE activates 104B parameters per token from a 2.8 trillion total pool. If you offload cold experts and keep only the hot routed experts in VRAM, the resident footprint drops to roughly 121 GB. The hybrid KDA and Gated-MLA attention keeps the key-value cache surprisingly small at full context.

By TechCompare · Updated

Total VRAM required
1781 GB
Kimi K3 2.8T (MoE) at Q4_K_M
Weights
1568 GB
2800B params
KV cache
51.5 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

1781 GB

2800B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
1568 GB
KV cache
51.5 GB
Overhead
162 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Hardware that fits

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

The 2.8T parameter pool requires 1568 GB of weights at Q4_K_M regardless of expert routing. Kimi K3's hybrid design uses 69 Kimi Delta Attention layers that store no key-value cache at all, plus 24 Gated-MLA layers that compress per-token KV to a 512-dim latent. Only the 24 MLA layers contribute to the cache, which lands at just 52 GB even at the full 1,048,576 token window. Activation and software overhead adds about 162 GB. In active-only mode the resident weights shrink to 58 GB, which drops total usage to 121 GB. The cache stays 52 GB either way because it depends only on context length, not expert loading.

Verdict

An all-resident Kimi K3 deployment is multi-node datacenter territory: roughly 19 H100 80GB or 11 H200 141GB with NVLink. Active-only fits on dual 80GB cards or a 128GB unified memory workstation, but cold-expert PCIe streaming limits throughput. The KDA plus MLA hybrid is what keeps the cache small enough for active-offload to even be viable at 1M context.

More Kimi scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
1600B - Q4_K_M - 1024K ctx
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
109B - Q4_K_M - 10240K ctx
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
20B - Q4_K_M - 128K ctx
View details ➜

Frequently asked questions

Why is Kimi K3's key-value cache so small for a 2.8T model?
Kimi K3 uses 69 Kimi Delta Attention layers that store no key-value cache at all, plus 24 Gated-MLA layers that compress per-token KV to a 512-dim latent. Only the 24 MLA layers contribute to the cache. At 1M context the cache is about 52 GB, far smaller than a comparable all-attention model of the same layer count.
Can I run Kimi K3 on a single GPU?
Not in resident mode. Even active-only needs 121 GB, which exceeds a single H100 80GB or H200 141GB. Active-only fits on a 128GB or 192GB unified memory workstation, but generation speeds will drop because cold experts stream over the PCIe bus.