TechCompare LogoTechCompare

How much VRAM does Llama 3.1 8B need at Q4_K_M? Single GPU local inference

If you have a working PC, you can run Llama 3.1 8B locally. Q4_K_M is the recommended quant - it fits everywhere and the quality loss vs FP16 is negligible at this scale. Step up to Q8_0 only if you have headroom and want maximum quality.

Llama 3.1 8B at Q4_K_M needs about 23.8 GB of VRAM at its native 128K context. This is what makes local inference of small models at high context highly attractive on 24 GB consumer GPUs.

By TechCompare · Updated

Total VRAM required
23.8 GB
Llama 3.1 8B at Q4_K_M
Weights
4.5 GB
8B params
KV cache
17.2 GB
128K tokens, FP16 KV

Calculator

Estimated VRAM required

23.8 GB

8B params at Q4_K_M, 131,072 token context, batch 1, inference.

Weights
4.5 GB
KV cache
17.2 GB
Overhead
2.2 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

KV cache exceeds model weights: Consider lowering the context length to save on VRAM. Contexts between 8K and 64K are generally more typical for local setups.

Hardware that fits

RTX 3090
Consumer
24 GB
99% used
RTX 5090
Consumer
32 GB
74% used
A100 40GB
Datacenter
40 GB
60% used
Apple M3 Max 64GB
Unified
48 GB
50% used

How this is calculated

8B at Q4_K_M is roughly 4.5 GB of weights plus 17.2 GB of KV cache and 2.1 GB of activation overhead, totaling 23.8 GB.

Verdict

The 23.8 GB total splits as 4.5 GB of Q4 weights, 17.2 GB of FP16 KV cache at the full 128K context (8 KV heads at head_dim 128 across 32 layers), and 2.1 GB of activation overhead. The headline insight is that the cache exceeds the weights at this context setting: the model itself is small, but the long-context budget isn't. Cap to 32K and the cache collapses to about 4.3 GB, dropping the total to about 11 GB, which fits any 12 GB consumer card with headroom. So the budget question isn't whether Llama 3.1 8B fits your hardware, it's how long your context is Q8 KV halves the cache and is a free quality lever when you need a few extra tokens in the budget.

More Llama scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
DeepSeek V4 Pro 1.6T at Q4_K_M with the full 1M-token context needs about 1012 GB of VRAM with every expert resident - this is the real shape of the model and the number to plan a deployment against.
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
Llama 4 Scout at Q4_K_M with its native 10M context needs about 2231 GB of VRAM with all 109B params resident - that's the number you size hardware against.
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
gpt-oss 20B at Q4_K_M with native 128K context needs about 19.4 GB of VRAM with all experts resident, dropping to roughly 9.3 GB with active-only weight loading.
View details ➜

Frequently asked questions

Can I run Llama 3.1 8B on a 6 GB GPU?
Yes at reduced context lengths. At the native 128K context, the 17.2 GB KV cache exceeds standard consumer limits, but dropping context to 8K brings VRAM down to 6.1 GB.
Is Llama 3.1 8B good enough for daily use?
For chat, summarization, and simple code completion, yes. For complex reasoning, math, or production-quality writing, you'll notice the gap to 70B-class models. Use 8B for speed and 70B+ for accuracy.
What's the cheapest GPU that runs Llama 3.1 8B at Q4?
A 6 GB GPU works. Q4 weights are ~4.3 GB, KV cache at 32K context adds roughly 1 GB, overhead is 0.6 GB, totaling about 6 GB. The RTX 2060 6 GB and GTX 1660 Super 6 GB run it. Modern 6 GB cards can also do small batches. For 8K context, an 8 GB card is the comfortable floor.
How much does context length change the VRAM requirement for Llama 3.1 8B?
Roughly linearly. The KV cache at 8K context is about 0.25 GB with FP16 KV, and at 128K it grows 16x to about 4 GB, rivaling the 4.3 GB weight footprint. That's why the same card that runs 8B comfortably at short context starts to struggle at the model's maximum window, and why Q8 KV cache quantization is the standard fix for long-context work.