How much VRAM does Qwen 3.8 2.4T Max need at Q4_K_M? Alibaba's biggest 1M MoE
Qwen 3.8 2.4T Max is a ~1839 GB resident deployment at 1M context, or ~419 GB with active-expert offload on four H200 cards. The bigger gate is the license: Qwen 3.8 Max carries non-commercial terms, so check usage rights before sizing hardware.
Qwen 3.8 2.4T Max at Q4_K_M with native 1M context needs about 1839 GB of VRAM with every expert resident. Alibaba's largest open release activates 95B parameters per token from the 2.4T pool. Active-expert offload drops the resident footprint to roughly 419 GB, but the non-commercial license terms mean most production deployments will reach for Alibaba's hosted API instead.
By TechCompare · Updated
Calculator
Estimated VRAM required
1856 GB
2400B params at Q4_K_M, 1,048,576 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
Hardware that fits
No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.
How this is calculated
The 2.4T parameter pool at Q4_K_M is 1344 GB of weights, the largest single number on this list after Kimi K3. The key-value cache with 8 KV heads at head_dim 128 across 80 layers is 328 GB at the full 1M context, and overhead adds about 167 GB. Active-only loading shrinks resident weights to 53.2 GB while the cache stays, totaling ~419 GB. That fits four 141 GB H200 cards, but the license caps commercial use at modest revenue thresholds.
Verdict
The 1839 GB resident budget splits into 1344 GB of weights, 328 GB of FP16 KV cache at 1M, and 167 GB of overhead - 23 H100 80GB cards or 13 H200 141GB cards resident. Active-only cuts weights to 53.2 GB and the total to 419 GB, which fits four H200s, but the per-token PCIe penalty of cold-expert fetch hurts throughput. The licensing reality matters more than the memory math for production: Qwen 3.8 Max ships with non-commercial terms, so for commercial traffic the Alibaba Cloud API or the fully Apache-2.0 Qwen3.5 line is the supported route.
More Qwen scenarios
Related guides
Frequently asked questions
Is Qwen 3.8 Max open for commercial use?
Should I run Qwen 3.8 2.4T Max or Qwen3.5 122B locally?
Related tools
RAM Latency Calculator
Convert DDR3/DDR4/DDR5 timings (CL, tRCD, tRP, tRAS) into true latency in nanoseconds.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜Memory and Storage Latency Visualizer
Visualize the massive speed difference between CPU cache, RAM, and storage.
Use tool ➜