TechCompare LogoTechCompare

DeepSeek V4.1 Flash VRAM: about 566 GB with Q4 backbone and FP8 Engram

This full-resident estimate needs about 566 GB: a Q4 552B backbone, FP8 196B Engram table with scales, and a logical FP16 cache at 1M tokens.

DeepSeek V4.1 Flash needs a weight budget that includes conditional memory as well as its MoE backbone. The publisher lists 552B backbone parameters and an additional 196B Engram table, for roughly 748B stored parameters in those two pools. Only about 8B parameters activate during prefill and 16B during decode. The preset keeps Engram at its reference FP8 storage precision and applies the selected weight quantization to the backbone.

By TechCompare · Updated

Total VRAM required
566 GB
DeepSeek V4.1 Flash 748B (MoE + Engram) at Q4_K_M backbone + FP8 Engram
Weights
511 GB
748B params
KV cache
3.4 GB
1024K tokens, FP16 KV

Calculator

Estimated VRAM required

566 GB

748B params at Q4_K_M backbone + FP8 Engram, 1,048,576 token context, batch 1, inference.

Weights
511 GB
KV cache
3.4 GB
Overhead
51.5 GB

Estimate accuracy: Planning estimate. Rounded model counts, tensor formats, cache packing, and workspace can change actual allocation. Confirm the exact checkpoint and serving engine with a load test.

DeepSeek V4.1 Flash includes a 552B backbone and 196B Engram table. All-resident weights keep Engram in FP8 with scales while the selected quant applies to the backbone. Cache estimates use logical FP16/Q8 storage, including shared cache owners and sliding buffers. Native packed FP4 cache, mixed backbone weights, and DSpark need a backend-specific budget. Custom weight budgets and estimates based on activated parameters use generic sizing and don't model Engram placement.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

At Q4_K_M's 0.56-byte planning factor, the 552B backbone costs 309.12 GB. The 196B Engram table remains FP8 with a one-byte scale per 32 values in the reference implementation, adding about 202.13 GB. That produces a 511.25 GB resident weight estimate before cache and runtime overhead. The 40-layer causal encoder-decoder shares global cache across four owner layers. Three encoder owners store a 512-value latent and 128-value index key per two tokens, while one decoder owner stores the same widths per token. Together they retain 1,600 logical values per input token. The preset also includes the backbone's 128-token sliding buffers and small FP32 compressor states. At 1,048,576 tokens, logical FP16 cache costs about 3.36 GB. Add 10% general overhead to reach 566.07 GB, about 527.19 GiB. This isn't the native packed-cache footprint: DeepSeek reports 890 bytes per token for its FP4 global cache. Native mixed backbone weights, draft buffers, and runtime packing require their own measurements.

Verdict

Engram storage dominates the difference between a backbone-only estimate and a complete resident budget. Shortening context from 1M to 32K lowers this modeled total only from 566.07 GB to 562.49 GB because shared cache is already compact. Moving Engram or experts to host memory could change VRAM substantially, but the table still needs storage and bandwidth somewhere. Use an implementation that explicitly supports that placement and measure lookup costs. Don't assume the active parameter count represents a deployable memory footprint.

More DeepSeek scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜

Frequently asked questions

Why does the preset say 748B when the model card says 552B?
The 552B figure describes the backbone. The same model card lists another 196B parameters in Engram conditional memory. The preset's total combines both rounded pools so an all-resident plan doesn't omit the lookup table. The weight calculation keeps their storage formats separate. Its 16B active figure refers to decode computation, while prefill activates about 8B according to the publisher.
Does selecting Q4 also quantize Engram to Q4?
No. In the full-resident preset, the selected format applies to the backbone while Engram stays FP8 with scale storage, matching the reference table representation. Assuming that every stored parameter uses the same Q4 factor would underbudget this component. Custom resident budgets and estimates based on activated parameters are generic scenarios and don't specify which Engram rows or expert tensors live on the GPU.
Why is the displayed FP16 cache larger than 890 bytes per token?
The 890-byte figure is DeepSeek's packed FP4 global-cache result with scale metadata. The tool's FP16 and Q8 choices estimate logical cache values at two or one byte each. They also include the reference backbone's fixed sliding buffers. These are different storage assumptions. The preset doesn't relabel native FP4 bytes as FP16, and it doesn't claim that every runtime supports the Q8 scenario.
Do all 40 layers keep a separate full context cache?
No. The reference configuration identifies four global KV owner layers. Intermediate layers reuse global KV or indexing information, and the decoder takes its global projection from the encoder output. The preset counts those shared owners instead of multiplying a full cache by 40. Local sliding buffers are still accounted for. Optional DSpark speculative paths have additional state that isn't itemized in this text-inference estimate.
What deployment information should I collect first?
Record the checkpoint revision, actual backbone tensor formats, Engram placement, cache packing, and speculative-decoding settings. Then measure resident memory and peak prefill memory on every device. A service can use less VRAM through host offload while requiring more system RAM and data movement. The memory calculator compares storage budgets, so it can't predict the latency of that design or certify compatibility with a serving framework.