How much VRAM does MiMo-V2.5 310B (MoE) need at Q4_K_M? Xiaomi 1M SWA model

MiMo-V2.5 310B at Q4_K_M is a 244 GB resident deployment for 1M context, or about 62 GB with active-expert offload.

MiMo-V2.5 310B at Q4_K_M with native 1M context needs about 244 GB of VRAM with all experts resident. The smaller sibling of MiMo-V2.5-Pro activates 15B parameters per token from a 310B total pool. Active-expert offload drops the resident footprint to roughly 62 GB. Sliding-window attention with a 128-token window keeps the key-value cache under 49 GB at the full 1M context.

By TechCompare · Updated

Total VRAM required
244 GB
MiMo-V2.5 310B (MoE) at Q4_K_M
Weights
174 GB
310B params
KV cache
48.3 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

244 GB

310B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
174 GB
KV cache
48.3 GB
Overhead
22.2 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Sliding-window attention applied: This model caps 4 of every 5 layers at a 128-token window. KV cache estimate is 81% smaller than naive full-attention math at this context length.

Hardware that fits

NVIDIA B300
Datacenter
288 GB
85% used

How this is calculated

The 310B parameter pool requires 174 GB of weights at Q4_K_M. The 48-layer architecture uses group-query attention with variable KV heads: 8 in the 9 full-attention layers, 4 in the 39 sliding-window layers. The QK and V dims share the same asymmetry as the Pro variant, so the effective head_dim is 160. Sliding-window attention caps 39 of 48 layers at a 128-token window. The key-value cache at the full 1M context uses about 48 GB. Activation and software overhead adds about 22 GB. In active-only mode the resident weights shrink to 8 GB, which drops total usage to 62 GB.

Verdict

Resident mode fits on four 80GB datacenter cards or two 141GB H200 cards. Active-only at 62 GB fits on a single 80GB card or a 96GB unified memory workstation. This is the practical option if you want the MiMo architecture on hand: the Pro variant needs 2.8x the VRAM for 2.8x the active parameters.

More MiMo scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
1600B - Q4_K_M - 1024K ctx
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
109B - Q4_K_M - 10240K ctx
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
20B - Q4_K_M - 128K ctx
View details ➜

Frequently asked questions

What's the difference between MiMo-V2.5 and MiMo-V2.5-Pro?
Same architecture family with asymmetric QK and V attention dims and sliding-window attention. The Pro variant has a 1.02T total pool with 42B active across 70 layers. The base variant has a 310B total pool with 15B active across 48 layers. The base is the practical pick if you have single-card or workstation hardware.
Does MiMo-V2.5 fit on a single GPU?
In active-only mode, yes. The 62 GB active-only total fits on an 80GB H100 or A100. Resident mode needs 244 GB, which requires four 80GB cards. Capping context to 256K drops the active-only total to about 30 GB, which fits on a single 40GB card.