How much VRAM does MiMo-V2.5-Pro 1.02T (MoE) need at Q4_K_M? Xiaomi 1M SWA frontier

MiMo-V2.5-Pro 1.02T at Q4_K_M is a 687 GB resident deployment for 1M context, or about 85 GB with active-expert offload.

MiMo-V2.5-Pro 1.02T at Q4_K_M with native 1M context needs about 687 GB of VRAM with all experts resident. Xiaomi's flagship MoE activates 42B parameters per token from a 1.02 trillion total pool. Active-expert offload drops the resident footprint to roughly 85 GB. The asymmetric QK and V attention dims plus sliding-window attention keep the key-value cache manageable even at full context.

By TechCompare · Updated

Total VRAM required
687 GB
MiMo-V2.5-Pro 1.02T (MoE) at Q4_K_M
Weights
571 GB
1020B params
KV cache
53.7 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

687 GB

1020B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
571 GB
KV cache
53.7 GB
Overhead
62.5 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Sliding-window attention applied: This model caps 6 of every 7 layers at a 128-token window. KV cache estimate is 86% smaller than naive full-attention math at this context length.

Hardware that fits

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

MiMo-V2.5-Pro's 1.02T parameter pool requires 571 GB of weights at Q4_K_M. The 70-layer architecture uses group-query attention with 8 KV heads and asymmetric key and value dims: keys use 192, values use 128, averaged to an effective head_dim of 160 in the cache math. Sliding-window attention caps 60 of 70 layers at a 128-token window with 10 layers doing full global attention. The result is a 54 GB key-value cache at the full 1M context. Activation and software overhead adds about 62 GB. In active-only mode the resident weights shrink to 24 GB, which drops total usage to 85 GB.

Verdict

Resident mode needs nine 80GB datacenter cards or five 141GB H200 cards. Active-only fits on a single 96GB pro card or a 128GB unified memory workstation. The sliding-window attention is why the cache stays under 55 GB even at 1M context. For most users the hosted Xiaomi API is the practical starting point.

More MiMo scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
1600B - Q4_K_M - 1024K ctx
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
109B - Q4_K_M - 10240K ctx
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
20B - Q4_K_M - 128K ctx
View details ➜

Frequently asked questions

Why does MiMo-V2.5-Pro use asymmetric QK and V dimensions?
Xiaomi splits the key and value projection dims. Keys use 192, values use 128. This trades a slightly larger key projection for a smaller value cache. The effective head_dim in the cache formula is the average of the two, 160. The result is a cache smaller than a symmetric 192-dim model but larger than a symmetric 128-dim model.
Can I run MiMo-V2.5-Pro locally?
Only in active-only mode. The 687 GB resident total needs multi-GPU datacenter hardware. Active-only at 85 GB fits on a 96 GB pro card or a high-memory unified workstation, but cold-expert PCIe streaming limits throughput. Capping context to 128K drops active-only to about 60 GB.