mistralai/Mistral-Small-4-119B-2603 VRAM requirements

119B params, Multi-head latent attention (MLA)

Mistral-Small-4-119B-2603 needs about 121 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 1× NVIDIA B200. It supports 64 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200164
NVIDIA H200 SXM264
NVIDIA H100 SXM264
NVIDIA H100 NVL264
NVIDIA A100 80GB SXM258
NVIDIA RTX PRO 6000 Blackwell Server Edition264
NVIDIA H20264
NVIDIA L40S464
32k context
GPUGPUsConcurrent users
NVIDIA B200164
NVIDIA H200 SXM264
NVIDIA H100 SXM227
NVIDIA H100 NVL260
NVIDIA A100 80GB SXM214
NVIDIA RTX PRO 6000 Blackwell Server Edition264
NVIDIA H20264
NVIDIA L40S423
128k context
GPUGPUsConcurrent users
NVIDIA B200126
NVIDIA H200 SXM238
NVIDIA H100 SXM26
NVIDIA H100 NVL215
NVIDIA A100 80GB SXM23
NVIDIA RTX PRO 6000 Blackwell Server Edition216
NVIDIA H20216
NVIDIA L40S45

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
FP8about 121 GB1x NVIDIA B200
INT4 (AWQ)about 67 GB1x NVIDIA B200

Model notes

Attention
Multi-head latent attention (MLA)
Architecture
Mixture-of-Experts
Layers
36
Hidden size
4,096
Vocabulary
131,072

MLA keeps one compact shared KV copy per GPU, so adding GPUs does not shrink the cache, but the copy itself is small for a model this size.

Frequently asked questions

Will mistralai/Mistral-Small-4-119B-2603 run on a single H100?
No, a single H100 SXM cannot hold mistralai/Mistral-Small-4-119B-2603 at 32k context with FP8 weights. The model requires multiple GPUs.
What is the cheapest GPU setup for mistralai/Mistral-Small-4-119B-2603?
The cheapest fitting setup at 32k context is 1x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does mistralai/Mistral-Small-4-119B-2603 need at 128k context?
At 128k context on the cheapest fitting setup, mistralai/Mistral-Small-4-119B-2603 uses about 121 GB of VRAM per GPU for FP8 weights and about 1.5 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run mistralai/Mistral-Small-4-119B-2603 with INT4 quantization?
Yes, mistralai/Mistral-Small-4-119B-2603 has INT4 weight figures. Its weights take about 67 GB in INT4 versus about 121 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →