TechCompare LogoTechCompare

MiMo-V2.6-Flash VRAM requirements: about 217 GB at Q4 and 1M context

MiMo-V2.6-Flash needs about 217 GB at Q4_K_M with 1,048,576 tokens and FP16 cache. At 32K context, the same model needs about 191 GB.

MiMo-V2.6-Flash stores about 309B parameters and activates about 15B per token. The smaller active count reduces computation, while the full weight pool still drives an all-resident deployment. This page estimates 173.04 GB for Q4 weights, 24.18 GB for the full-context text cache, and 19.72 GB of runtime allowance. The resulting 216.95 GB is a planning estimate for one sequence, with the backbone's mixed attention layout included.

By TechCompare · Updated

Total VRAM required
217 GB
MiMo-V2.6-Flash 309B (MoE) at Q4_K_M
Weights
173 GB
309B params
KV cache
24.2 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

217 GB

309B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
173 GB
KV cache
24.2 GB
Overhead
19.7 GB

Estimate accuracy: Planning estimate. Rounded model counts, tensor formats, cache packing, and workspace can change actual allocation. Confirm the exact checkpoint and serving engine with a load test.

MiMo-V2.6-Flash has 9 global layers with 4 KV heads and 39 sliding layers with 8 KV heads. Keys are 192 wide and values are 128 wide. Vision/audio processing and the optional speculative decoder need extra runtime memory beyond this text-cache budget.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

Apple M3 Ultra 256GB
Unified
~248 GB
87% of estimate
NVIDIA B300
Datacenter
288 GB
75% used

Just barely too small

MI300X
Datacenter
192 GB
short by 24.9 GB

How this is calculated

The published Flash configuration has 48 layers and hidden size 4,096. Nine layers use global attention with four KV heads. The other 39 use a 128-token sliding window with eight KV heads. Both use 192-wide keys and 128-wide values. Using one head count for the entire model would size one of those groups incorrectly, so the preset tracks their widths separately. The global cache grows with context, while local buffers stop at the window size. At 32,768 tokens their combined FP16 cache is about 0.78 GB, producing a 191.20 GB total. At 1,048,576 tokens the total reaches 216.95 GB, or about 202.05 GiB. Q4_K_M uses an approximate 0.56-byte weight factor. That scenario doesn't reproduce the exact precision mix of Xiaomi's release, and the general runtime allowance doesn't separately model image/audio activations or the optional five-layer speculative decoder.

Verdict

MiMo Flash is much easier to place in memory than Pro, but its all-resident weights still exceed ordinary single-GPU capacity. Three nominal 80 GB GPUs offer 240 GB in aggregate, which exceeds this estimate on paper. That comparison doesn't guarantee a valid three-way tensor split, enough memory on each shard, or a compatible serving kernel. Select the runtime and checkpoint first, inspect its parallelism constraints, and measure the workload before treating the aggregate memory figure as a deployment plan.

More MiMo scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜
DeepSeek V4.1 Flash Q4 + FP8 Engram
Include 552B backbone weights and the separate 196B Engram table.
View details ➜

Frequently asked questions

Is MiMo Flash a 15B model for VRAM purposes?
Its 15B activated parameter count is a compute figure. Full-speed all-resident inference has to hold the roughly 309B pool because later tokens can route to other experts. CPU or storage offload can move part of that pool out of VRAM, but it changes the system's memory and bandwidth requirements. Don't substitute the active count for a measured resident tensor budget.
Why does this preset show four KV heads?
Four is the global-attention head count from the configuration. The preset also records eight KV heads for the sliding layers and uses that second value in the calculation. The 160 head dimension displayed by the generic tool is an effective average of a 192-wide key and 128-wide value. It preserves the combined cache width without pretending the two tensors have identical dimensions.
Can I save memory by lowering context?
Yes. Moving from 1M to 32K reduces this modeled total by about 25.74 GB. The global cache shrinks while the short sliding buffers remain bounded. The 173.04 GB Q4 weight budget stays the same, so smaller context doesn't turn Flash into a small local model. A supported lower cache precision can reduce the sequence memory further without changing the expert pool.
Does the article cover multimodal and speculative serving?
It covers the backbone weight planning factor and text attention cache. Xiaomi also publishes vision/audio components and a five-layer speculative decoder with its own local attention. Those paths need separate activation and draft-buffer allowances when enabled. The rounded model count and 10% overhead shouldn't be treated as an exact accounting of every component. Check the runtime's peak allocation with representative multimodal requests.
Should I use the RL or MOPD checkpoint for this estimate?
The published backbone architecture is shared, so the preset can size either at this level of detail. The linked MOPD model card and configuration identify the reference used here. Choose a checkpoint based on task behavior and serving support, then confirm its tensor formats. A generic Q4 budget says how much memory that precision might require, not whether a specific conversion or inference engine already supports it.