MiniMaxAI/Minimax-M3 VRAM requirements

427B params, Grouped-query attention (GQA)

Minimax-M3 needs about 429 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 4× NVIDIA B200. It supports 64 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200464
NVIDIA H200 SXM464
NVIDIA H100 SXM864
NVIDIA H100 NVL864
NVIDIA A100 80GB SXM864
NVIDIA RTX PRO 6000 Blackwell Server Edition864
NVIDIA H20864
NVIDIA L40S16 (verify server configuration)64
32k context
GPUGPUsConcurrent users
NVIDIA B200464
NVIDIA H200 SXM420
NVIDIA H100 SXM833
NVIDIA H100 NVL858
NVIDIA A100 80GB SXM823
NVIDIA RTX PRO 6000 Blackwell Server Edition861
NVIDIA H20861
NVIDIA L40S16 (verify server configuration)44
128k context
GPUGPUsConcurrent users
NVIDIA B200426
NVIDIA H200 SXM45
NVIDIA H100 SXM88
NVIDIA H100 NVL814
NVIDIA A100 80GB SXM85
NVIDIA RTX PRO 6000 Blackwell Server Edition815
NVIDIA H20815
NVIDIA L40S16 (verify server configuration)11

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
BF16about 854 GB8x NVIDIA B200
FP8about 429 GB4x NVIDIA B200
INT4 (AWQ)about 237 GB2x NVIDIA B200

Model notes

Attention
Grouped-query attention (GQA)
Architecture
Mixture-of-Experts
Layers
60
Hidden size
6,144
Vocabulary
200,064
Native precision
bfloat16

At 32k context, 32% of this card's usable memory is working as conversation cache.

Frequently asked questions

Will MiniMaxAI/Minimax-M3 run on a single H100?
No, a single H100 SXM cannot hold MiniMaxAI/Minimax-M3 at 32k context with FP8 weights. The model requires multiple GPUs.
What is the cheapest GPU setup for MiniMaxAI/Minimax-M3?
The cheapest fitting setup at 32k context is 4x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does MiniMaxAI/Minimax-M3 need at 128k context?
At 128k context on the cheapest fitting setup, MiniMaxAI/Minimax-M3 uses about 107 GB of VRAM per GPU for FP8 weights and about 2.0 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run MiniMaxAI/Minimax-M3 with INT4 quantization?
Yes, MiniMaxAI/Minimax-M3 has INT4 weight figures. Its weights take about 237 GB in INT4 versus about 429 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →