TechCompare LogoTechCompare

How much VRAM does Kimi K3 2.8T (MoE) need at Q4_K_M? Moonshot 1M frontier

Kimi K3 2.8T at Q4_K_M is a roughly 1.8 TB resident deployment for 1M context, or about 121 GB if you can tolerate active-expert offload latency.

Kimi K3 2.8T at Q4_K_M with native 1M context needs about 1781 GB of VRAM for an all-resident deployment. Moonshot's third-generation trillion-parameter MoE activates 104B parameters per token from a 2.8 trillion total pool. If you offload cold experts and keep only the hot routed experts in VRAM, the resident footprint drops to roughly 121 GB. The hybrid KDA and Gated-MLA attention keeps the key-value cache surprisingly small at full context.

By TechCompare · Updated

Total VRAM required
1781 GB
Kimi K3 2.8T (MoE) at Q4_K_M
Weights
1568 GB
2800B params
KV cache
51.5 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

1781 GB

2800B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
1568 GB
KV cache
51.5 GB
Overhead
162 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

The 2.8T parameter pool requires 1568 GB of weights at Q4_K_M regardless of expert routing. Kimi K3's hybrid design uses 69 Kimi Delta Attention layers that store no key-value cache at all, plus 24 Gated-MLA layers that compress per-token KV to a 512-dim latent. Only the 24 MLA layers contribute to the cache, which lands at just 52 GB even at the full 1,048,576 token window. Activation and software overhead adds about 162 GB. In active-only mode the resident weights shrink to 58 GB, which drops total usage to 121 GB. The cache stays 52 GB either way because it depends only on context length, not expert loading.

Verdict

An all-resident Kimi K3 deployment is multi-node datacenter territory: roughly 19 H100 80GB or 11 H200 141GB with NVLink. Active-only fits on dual 80GB cards or a 128GB unified memory workstation, but cold-expert PCIe streaming limits throughput. The KDA plus MLA hybrid is what keeps the cache small enough for active-offload to even be viable at 1M context.

More Kimi scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
DeepSeek V4 Pro 1.6T at Q4_K_M with the full 1M-token context needs about 1012 GB of VRAM with every expert resident - this is the real shape of the model and the number to plan a deployment against.
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
At Q4_K_M and 10,485,760 tokens, this model estimates 2335 GB with all 109B parameters resident.
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
gpt-oss 20B at Q4_K_M with native 128K context needs about 19.4 GB of VRAM with all experts resident, dropping to roughly 9.3 GB with active-only weight loading.
View details ➜

Frequently asked questions

Why is Kimi K3's key-value cache so small for a 2.8T model?
Kimi K3 uses 69 Kimi Delta Attention layers that store no key-value cache at all, plus 24 Gated-MLA layers that compress per-token KV to a 512-dim latent. Only the 24 MLA layers contribute to the cache. At 1M context the cache is about 52 GB, far smaller than a comparable all-attention model of the same layer count.
Can I run Kimi K3 on a single GPU?
Not in resident mode. Even active-only needs 121 GB, which exceeds a single H100 80GB or H200 141GB. Active-only fits on a 128GB or 192GB unified memory workstation, but generation speeds will drop because cold experts stream over the PCIe bus.
Can I run Kimi K3 2.8T on a single GPU?
No. The Q4 weights alone are roughly 1.57 TB. Even with active-only loading (where the resident expert subset fits in 60 GB), the cold-expert fetch over PCIe bottleneck makes per-token throughput impractical for production. Kimi K3's value prop is its compressed MLA key-value cache, but the parameter pool is still too large for consumer GPUs.
What is MLA and why does Kimi K3's KV cache matter?
Multi-head Latent Attention compresses the KV cache by projecting key and value vectors into a smaller latent space before storing them, which multiplies down the VRAM cost of long context. It's the same family of trick DeepSeek uses, and it's Kimi K3's main memory story. The weight pool still dominates the bill at 2.8T total parameters, but MLA is what makes the long-context math plausible at all.