TechCompare LogoTechCompare

How much VRAM does Qwen 3.8 27B need at Q4_K_M? Dense 1M context mid-size

Qwen 3.8 27B at Q4 needs ~45 GB at 128K context or ~244 GB at the full native 1M window - the KV cache, not the 15 GB of weights, sets the hardware. A 48 GB card handles it at 128K, and that's how most people should run it.

Qwen 3.8 27B at Q4_K_M needs about 244 GB of VRAM if you stretch it to the full native 1M context, but the weights are only 15 GB - the rest is key-value cache. Alibaba's dense 27B is the practical local-deployable member of the Qwen 3.8 family, with over a million downloads on HuggingFace since release.

By TechCompare · Updated

Total VRAM required
244 GB
Qwen 3.8 27B at Q4_K_M
Weights
15.4 GB
27.5B params
KV cache
206 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

244 GB

27.5B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
15.4 GB
KV cache
206 GB
Overhead
22.2 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

KV cache exceeds model weights: Consider lowering the context length to save on VRAM. Contexts between 8K and 64K are generally more typical for local setups.

Hardware that fits

NVIDIA B300
Datacenter
288 GB
85% used

How this is calculated

Weights at Q4_K_M are 15.4 GB. The KV cache with 8 KV heads at head_dim 128 across 48 layers is 206 GB at the full 1M context, 25.8 GB at 128K, and overhead adds 22 GB at the long window. This is the dense-model long-context trade in the open: a mid-size model with a 1M native window pays for that window in cache, not weights. At the realistic 128K working context, total drops to about 45 GB and the model fits on a 48 GB card.

Verdict

The number to plan around is 45 GB at 128K context: 15.4 GB of Q4 weights, 25.8 GB of FP16 KV cache, 4.1 GB of overhead. That fits a 48 GB RTX 6000 Ada or two 24 GB consumer cards. Stretching to the native 1M window inflates the cache to 206 GB and the total to 244 GB, which needs datacenter RAM or serious GPU pooling for a dense model this small. Most workloads cap out under 128K, and RAG beats raw long-context for retrieval-shaped tasks anyway, so run it at 128K and let the 1M number stand for the worst case.

More Qwen scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
DeepSeek V4 Pro 1.6T at Q4_K_M with the full 1M-token context needs about 1012 GB of VRAM with every expert resident - this is the real shape of the model and the number to plan a deployment against.
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
Llama 4 Scout at Q4_K_M with its native 10M context needs about 2231 GB of VRAM with all 109B params resident - that's the number you size hardware against.
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
gpt-oss 20B at Q4_K_M with native 128K context needs about 19.4 GB of VRAM with all experts resident, dropping to roughly 9.3 GB with active-only weight loading.
View details ➜

Frequently asked questions

Why does a 27B model need 233 GB at 1M context?
The KV cache scales with context length times layers times KV width, not with weights. At 1M tokens across 48 layers with 8 KV heads at head_dim 128, the cache is 206 GB - over 13x the size of the Q4 weights. At 128K the same math yields 25.8 GB and the model fits mid-range pro cards.
Qwen 3.8 27B or Qwen3.5 27B?
Qwen 3.8 27B if you need the 1M native window and the newer training data. Qwen3.5 27B has nearly identical VRAM math (same family scaling) at a 256K native context. For RAG and standard agentic workloads at 128K or below, they're equivalent on hardware.