moonshotai/Kimi-K2.6 VRAM requirements

1059B params, Multi-head latent attention (MLA)

Kimi-K2.6 needs about 595 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 4× NVIDIA B200. It supports 9 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200439
NVIDIA H200 SXM864
NVIDIA H100 NVL829
NVIDIA RTX PRO 6000 Blackwell Server Edition836
NVIDIA H20836
NVIDIA H100 SXM16 (verify server configuration)64
NVIDIA A100 80GB SXM16 (verify server configuration)64
NVIDIA L40S16 (verify server configuration)2
32k context
GPUGPUsConcurrent users
NVIDIA B20049
NVIDIA H200 SXM837
NVIDIA H100 NVL87
NVIDIA RTX PRO 6000 Blackwell Server Edition89
NVIDIA H2089
NVIDIA H100 SXM16 (verify server configuration)54
NVIDIA A100 80GB SXM16 (verify server configuration)46
NVIDIA L40Sno fit
128k context
GPUGPUsConcurrent users
NVIDIA B20042
NVIDIA H200 SXM89
NVIDIA H100 NVL81
NVIDIA RTX PRO 6000 Blackwell Server Edition82
NVIDIA H2082
NVIDIA H100 SXM16 (verify server configuration)13
NVIDIA A100 80GB SXM16 (verify server configuration)11
NVIDIA L40Sno fit

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
FP8about 595 GB4x NVIDIA B200
INT4 (AWQ)about 582 GB4x NVIDIA B200

Model notes

Attention
Multi-head latent attention (MLA)
Architecture
Mixture-of-Experts
Layers
61
Hidden size
7,168
Vocabulary
163,840

MLA keeps one compact shared KV copy per GPU, so adding GPUs does not shrink the cache, but the copy itself is small for a model this size.

Frequently asked questions

Will moonshotai/Kimi-K2.6 run on a single H100?
No, a single H100 SXM cannot hold moonshotai/Kimi-K2.6 at 32k context with FP8 weights. The model requires multiple GPUs.
What is the cheapest GPU setup for moonshotai/Kimi-K2.6?
The cheapest fitting setup at 32k context is 4x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does moonshotai/Kimi-K2.6 need at 128k context?
At 128k context on the cheapest fitting setup, moonshotai/Kimi-K2.6 uses about 149 GB of VRAM per GPU for FP8 weights and about 4.6 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run moonshotai/Kimi-K2.6 with INT4 quantization?
Yes, moonshotai/Kimi-K2.6 has INT4 weight figures. Its weights take about 582 GB in INT4 versus about 595 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →