Estimated VRAM required
15.0 GB
9B params at Q4_K_M, 262,144 token context, batch 1, inference.
Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA or hybrid state estimates. Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.
KV cache exceeds model weights: Consider lowering the context length to save on VRAM. Contexts between 8K and 64K are generally more typical for local setups.
Hybrid linear attention applied: 24 of 32 layers use fixed recurrent state, so only 8 full-attention layers contribute to the context-growing KV cache. The fixed state adds about 0.05 GB at batch 1 and does not grow with context length.
Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.