TechCompare LogoTechCompare

How much VRAM does Qwen 3.8 2.4T Max need at Q4_K_M? Alibaba's biggest 1M MoE

Qwen 3.8 2.4T Max is a ~1839 GB resident deployment at 1M context. The bigger gate is the license: Qwen 3.8 Max carries non-commercial terms, so check usage rights before sizing hardware.

Qwen 3.8 2.4T Max at Q4_K_M with native 1M context needs about 1839 GB of VRAM with every expert resident. Alibaba's largest open release activates 95B parameters per token from the 2.4T pool. That doesn't shrink what has to stay loaded, because different tokens pick different experts. The non-commercial license terms mean most production deployments will reach for Alibaba's hosted API instead.

By TechCompare · Updated

Total VRAM required
1856 GB
Qwen 3.8 2.4T Max (MoE) at Q4_K_M
Weights
1344 GB
2400B params
KV cache
344 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

1856 GB

2400B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
1344 GB
KV cache
344 GB
Overhead
169 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

The 2.4T parameter pool at Q4_K_M is 1344 GB of weights, the largest single number on this list after Kimi K3. The key-value cache with 8 KV heads at head_dim 128 across 80 layers is 328 GB at the full 1M context, and overhead adds about 167 GB. The license caps commercial use at modest revenue thresholds.

Verdict

The 1839 GB resident budget splits into 1344 GB of weights, 328 GB of FP16 KV cache at 1M, and 167 GB of overhead - 23 H100 80GB cards or 13 H200 141GB cards resident. The 95B active count doesn't lower the resident total, because any token can route to any expert. The licensing reality matters more than the memory math for production: Qwen 3.8 Max ships with non-commercial terms, so for commercial traffic the Alibaba Cloud API or the fully Apache-2.0 Qwen3.5 line is the supported route.

More Qwen scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜

Frequently asked questions

Is Qwen 3.8 Max open for commercial use?
Not fully. Qwen 3.8 Max ships with revenue-share style restrictions that cap commercial deployment, unlike the Apache-2.0 Qwen3.5 line. For production commercial workloads, use the hosted Alibaba Cloud API or stick with Qwen3.5 122B, which is fully permissive.
Should I run Qwen 3.8 2.4T Max or Qwen3.5 122B locally?
Qwen3.5 122B for almost everyone. It fits on a single H200 141GB resident and has clean Apache-2.0 licensing. The 2.4T Max needs 13+ H200s resident and carries license limitations that matter the moment you ship to customers.