TechCompare LogoTechCompare

K2 Horizon 375B A23B VRAM: about 375 GB at Q4 and 512K context

K2 Horizon 375B A23B needs about 375 GB with Q4_K_M weights, FP16 cache, and 524,288 tokens. At 32K context, the estimate is about 240 GB.

K2 Horizon 375B A23B is IFM's MoE model with roughly 375B stored parameters and 23B active per token. Its native 524,288-token context makes cache memory a large part of the full-context budget. At Q4_K_M, weights take about 210 GB and the FP16 attention cache takes another 131 GB. A 10% runtime allowance brings the total to 375.10 GB for one sequence. The active count doesn't replace the stored expert pool.

By TechCompare · Updated

Total VRAM required
375 GB
K2 Horizon 375B A23B (MoE) at Q4_K_M
Weights
210 GB
375B params
KV cache
131 GB
512K tokens, FP16 KV

Calculator

Estimated VRAM required

375 GB

375B params at Q4_K_M, 524,288 token context, batch 1, inference.

Weights
210 GB
KV cache
131 GB
Overhead
34.1 GB

Estimate accuracy: Planning estimate. Rounded model counts, tensor formats, cache packing, and workspace can change actual allocation. Confirm the exact checkpoint and serving engine with a load test.

K2 Horizon uses 61 full-context GQA layers with 8 KV heads. Its 23B active count describes computation, while all-resident weights include the full 375B pool. Quantization choices are planning scenarios and require a compatible checkpoint and backend.

Hardware that fits

Apple entries show installed unified memory separately from a model budget based on an 8 GB reserve. That budget is a planning heuristic, not a fixed hardware limit.

No single GPU in our catalog has enough memory. Multi-GPU or CPU offload required.

How this is calculated

The released configuration specifies 61 layers, hidden size 6,144, 48 query heads, eight KV heads, and a 128-value head dimension. Grouped-query attention shares cached keys and values across query heads, so memory uses the eight KV heads. This preset has no sliding-window or shared-cache reduction. The formula is two tensors times 61 layers times eight KV heads times 128 values times context times bytes per value. At 524,288 tokens in FP16 it yields 130.9965 GB. Add 210 GB for approximate Q4 weights and 34.10 GB of overhead to reach 375.10 GB, or about 349.34 GiB. At 32,768 tokens the cache falls to 8.19 GB and the total to 240.01 GB. These are inference budgets using rounded published parameter counts. The Q4 factor doesn't establish that a compatible converted checkpoint exists, and the runtime allowance doesn't explicitly simulate prompt activations or communication buffers.

Verdict

Choose context and concurrency before selecting the memory pool. Moving from 32K to the full 512K window adds about 135 GB to this modeled total, enough to change a multi-GPU plan. Five nominal 80 GB GPUs exceed 375 GB in aggregate, but that says nothing about whether a five-way split is supported or balanced. Prefer a validated parallelism layout with measured per-device headroom. For many shorter requests, a smaller context cap may be more useful than reserving the entire native window for every sequence.

More K2 Horizon scenarios

MiMo-V2.6-Pro at Q4_K_M
Budget a 1.02T expert pool, hybrid attention, and a 1M text context.
View details ➜
MiMo-V2.6-Flash at Q4_K_M
309B total weights and separate global and sliding cache widths.
View details ➜
GLM-5.3-Flash at Q4_K_M
Model the 34 KDA layers, 11 sparse layers, and indexer storage.
View details ➜

Frequently asked questions

Does A23B mean I need memory for only 23B parameters?
A23B is the active parameter count per token. The full model still stores about 375B parameters, and routing can select different experts for different tokens. All-resident inference therefore budgets the complete pool. Offloading can reduce the GPU-resident portion, but the rest needs host memory or another storage tier and a serving strategy that handles the resulting data movement.
Why is the cache so large despite GQA?
Eight KV heads are fewer than the 48 query heads, which already saves substantial memory compared with caching all query heads. There are still 61 layers and up to 524,288 token positions per sequence. Multiplying those dimensions makes the full FP16 cache about 131 GB. GQA reduces cache width, but it doesn't stop the cache from growing with context.
What happens if I select Q8 cache?
The calculator uses one byte per logical cache value rather than two, so this cache budget falls from about 131 GB to 65.50 GB at full context. With the same Q4 weights and 10% allowance, the total becomes about 303.05 GB. Treat that as a memory scenario until your serving engine confirms compatible cache quantization and its actual scale, metadata, and alignment overhead.
How should I budget concurrent requests?
Weights can be shared across requests, but each sequence has its own cached token positions. With two full 512K sequences, the modeled FP16 cache is about 262 GB rather than 131 GB. Real continuous-batching systems allocate cache pages according to their scheduler, so set context and concurrency limits together. Also test prompt-processing peaks, which can exceed a steady decode allocation.
Is a GPU with enough aggregate memory guaranteed to work?
No. You need support for the model's architecture and tensor formats, a valid parallelism split, and enough memory on each participating device. KV cache and other tensors may be replicated rather than evenly divided, depending on the engine. Use the linked publisher configuration to check dimensions, then validate the exact runtime setup. The calculator reports a storage budget and doesn't predict tokens per second.