TechCompare LogoTechCompare

How much VRAM does Mistral Medium 3.5 128B need at Q4_K_M? Mistral's open flagship

Mistral Medium 3.5 is a two-card model at Q4: ~150 GB at 256K context, or ~114 GB at 128K on a single 141 GB H200. Dense means no offload cheats, but also no per-token routing penalties.

Mistral Medium 3.5 128B at Q4_K_M needs about 150 GB of VRAM at its native 256K context, so a single 192 GB MI300X or two 80 GB cards run it fully resident. Mistral's July 2026 open-weights release finally gives the Medium line a downloadable model, dense at 128B with no expert routing.

By TechCompare · Updated

Total VRAM required
150 GB
Mistral Medium 3.5 128B at Q4_K_M
Weights
71.7 GB
128B params
KV cache
64.4 GB
256K tokens, FP16 KV

Calculator

Estimated VRAM required

150 GB

128B params at Q4_K_M, 262,144 token context, batch 1, inference.

Weights
71.7 GB
KV cache
64.4 GB
Overhead
13.6 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

MI300X
Datacenter
192 GB
78% used
Apple M3 Ultra 256GB
Unified
~248 GB
60% of estimate
NVIDIA B300
Datacenter
288 GB
52% used

Just barely too small

H200 141GB
Datacenter
141 GB
short by 8.7 GB

How this is calculated

Weights at Q4_K_M are 71.7 GB. The KV cache with 8 KV heads at head_dim 128 across 60 layers is 64.4 GB at the full 256K window, plus 13.6 GB of overhead. Being dense, there's no expert-offload trick - every parameter must stay resident, so the floor is set at 256K context by the cache. Capping context to 128K halves the cache to 32 GB and the total to ~114 GB, which fits a single 141 GB H200.

Verdict

The 150 GB total at 256K comes from 71.7 GB of Q4 weights plus 64.4 GB of FP16 KV cache plus 13.6 GB of overhead. Dense architecture means there are no expert weights to move to system RAM, but inference is also simpler, with predictable throughput. Versus Qwen3.5 122B at the same footprint class, Medium 3.5 trades a slightly larger cache at 256K for Mistral's agentic tuning and European data-residency story. A single MI300X 192GB or H200 pair handles it at full window.

More Mistral scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜

Frequently asked questions

Is Mistral Medium 3.5 open weights?
Yes - the 3.5 release finally made the Medium line downloadable, under Mistral's standard open license. Earlier Medium versions were API-only.
Mistral Medium 3.5 or Qwen3.5 122B?
Both land in the ~100-150 GB class at Q4. Qwen3.5 122B is MoE with a 104 GB footprint, and its expert weights can move to system RAM when VRAM runs short, at a real speed cost. Medium 3.5 is dense and simpler to serve with predictable per-token latency, and the Mistral stack integrates more cleanly with European hosting stacks. Benchmark both on your workload - they're close enough that ecosystem decides.