How much VRAM does Inkling 975B (MoE) need at Q4_K_M? Thinking Machines 1M hybrid

Inkling 975B at Q4_K_M is a 653 GB resident deployment for 1M context, or about 77 GB with active-expert offload.

Inkling 975B at Q4_K_M with native 1M context needs about 653 GB of VRAM with all experts resident. Thinking Machines' flagship MoE activates 41B parameters per token from a 975B total pool. The hybrid global and sliding-window attention keeps the key-value cache under 48 GB at the full 1M context. Active-expert offload drops the resident footprint to roughly 77 GB.

By TechCompare · Updated

Total VRAM required
653 GB
Inkling 975B (MoE) at Q4_K_M
Weights
546 GB
975B params
KV cache
47.4 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

653 GB

975B params at Q4_K_M, 1,048,576 token context, batch 1, inference.

Weights
546 GB
KV cache
47.4 GB
Overhead
59.3 GB

Estimate accuracy: Weights within ~2%. KV cache within ~5% for standard GQA models, ~10% for MLA (DeepSeek). Real VRAM may vary with framework (vLLM vs llama.cpp vs Transformers), Flash Attention, and driver overhead.

Sliding-window attention applied: This model caps 5 of every 6 layers at a 512-token window. KV cache estimate is 83% smaller than naive full-attention math at this context length.

Hardware that fits

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

The 975B parameter pool requires 546 GB of weights at Q4_K_M. The 66-layer architecture uses group-query attention with variable KV heads: 8 in the 11 global layers, 16 in the 55 sliding-window layers. The sliding window caps 55 of 66 layers at a 512-token window with only 11 layers doing full global attention. The result is a 47 GB key-value cache at the full 1M context. Activation and software overhead adds about 59 GB. In active-only mode the resident weights shrink to 23 GB, which drops total usage to 77 GB.

Verdict

Resident mode needs nine 80GB datacenter cards or five 141GB H200 cards. Active-only at 77 GB fits on a single 80GB card with headroom or a 96GB unified memory workstation. The 512-token sliding window on most layers is what keeps the cache manageable at full context. For most workloads the hosted API is the practical starting point.

More Inkling scenarios

DeepSeek V4 Pro 1.6T (MoE) at Q4_K_M
1600B - Q4_K_M - 1024K ctx
View details ➜
Llama 4 Scout (17B/109B) at Q4_K_M
109B - Q4_K_M - 10240K ctx
View details ➜
gpt-oss 20B (MoE) at Q4_K_M
20B - Q4_K_M - 128K ctx
View details ➜

Frequently asked questions

Why does Inkling use variable KV heads across layers?
The 11 global-attention layers use 8 KV heads while the 55 sliding-window layers use 16 KV heads. The sliding-window layers are capped at a 512-token window, so their extra KV heads cost little at long context. The global layers dominate the cache at full context, so the effective count used in sizing is 8.
Can I run Inkling on a single GPU?
In active-only mode, yes. The 77 GB active-only total fits on an 80GB H100 or A100. Resident mode needs 653 GB across nine 80GB cards. Capping context to 256K drops active-only to about 54 GB, which leaves comfortable headroom on a single 80GB card.