deepseek-ai/DeepSeek-V4-Flash VRAM requirements

158B params, Compressed recurrent state

DeepSeek-V4-Flash needs about 156 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 1× NVIDIA B200. It supports 3 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200111
NVIDIA H200 SXM264
NVIDIA H100 NVL216
NVIDIA RTX PRO 6000 Blackwell Server Edition222
NVIDIA H20222
NVIDIA H100 SXM464
NVIDIA A100 80GB SXM464
NVIDIA L40S862
32k context
GPUGPUsConcurrent users
NVIDIA B20013
NVIDIA H200 SXM234
NVIDIA H100 NVL24
NVIDIA RTX PRO 6000 Blackwell Server Edition25
NVIDIA H2025
NVIDIA H100 SXM426
NVIDIA A100 80GB SXM422
NVIDIA L40S816
128k context
GPUGPUsConcurrent users
NVIDIA B200217
NVIDIA H200 SXM28
NVIDIA H100 NVL21
NVIDIA RTX PRO 6000 Blackwell Server Edition21
NVIDIA H2021
NVIDIA H100 SXM46
NVIDIA A100 80GB SXM45
NVIDIA L40S84

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
FP8about 156 GB1x NVIDIA B200
INT4 (AWQ)about 86 GB1x NVIDIA B200

Model notes

Attention
Compressed recurrent state
Architecture
Mixture-of-Experts
Layers
43
Hidden size
4,096
Vocabulary
129,280
Native precision
bfloat16

This model keeps a fixed-size recurrent state per conversation instead of a growing token cache, so very long chats cost little extra memory.

Frequently asked questions

Will deepseek-ai/DeepSeek-V4-Flash run on a single H100?
No, a single H100 SXM cannot hold deepseek-ai/DeepSeek-V4-Flash at 32k context with FP8 weights. The model requires multiple GPUs.
What is the cheapest GPU setup for deepseek-ai/DeepSeek-V4-Flash?
The cheapest fitting setup at 32k context is 1x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does deepseek-ai/DeepSeek-V4-Flash need at 128k context?
At 128k context on the cheapest fitting setup, deepseek-ai/DeepSeek-V4-Flash uses about 156 GB of VRAM per GPU for FP8 weights and about 4.6 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run deepseek-ai/DeepSeek-V4-Flash with INT4 quantization?
Yes, deepseek-ai/DeepSeek-V4-Flash has INT4 weight figures. Its weights take about 86 GB in INT4 versus about 156 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →