Qwen/Qwen3-235B-A22B VRAM requirements

235B params, Grouped-query attention (GQA)

Qwen3-235B-A22B needs about 236 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 2× NVIDIA B200. It supports 26 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200264
NVIDIA H200 SXM464
NVIDIA H100 SXM458
NVIDIA H100 NVL464
NVIDIA A100 80GB SXM434
NVIDIA RTX PRO 6000 Blackwell Server Edition464
NVIDIA H20464
NVIDIA L40S846
32k context
GPUGPUsConcurrent users
NVIDIA B200226
NVIDIA H200 SXM464
NVIDIA H100 SXM414
NVIDIA H100 NVL430
NVIDIA A100 80GB SXM48
NVIDIA RTX PRO 6000 Blackwell Server Edition433
NVIDIA H20433
NVIDIA L40S811
128k context
GPUGPUsConcurrent users
NVIDIA B20026
NVIDIA H200 SXM418
NVIDIA H100 SXM43
NVIDIA H100 NVL47
NVIDIA A100 80GB SXM42
NVIDIA RTX PRO 6000 Blackwell Server Edition48
NVIDIA H2048
NVIDIA L40S82

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
BF16about 470 GB4x NVIDIA B200
FP8about 236 GB2x NVIDIA B200
INT4 (AWQ)about 130 GB1x NVIDIA B200

Model notes

Attention
Grouped-query attention (GQA)
Architecture
Mixture-of-Experts
Layers
94
Hidden size
4,096
Vocabulary
151,936
Native precision
bfloat16

At 32k context, 26% of this card's usable memory is working as conversation cache.

Frequently asked questions

Will Qwen/Qwen3-235B-A22B run on a single H100?
No, a single H100 SXM cannot hold Qwen/Qwen3-235B-A22B at 32k context with FP8 weights. The model requires multiple GPUs.
What is the cheapest GPU setup for Qwen/Qwen3-235B-A22B?
The cheapest fitting setup at 32k context is 2x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does Qwen/Qwen3-235B-A22B need at 128k context?
At 128k context on the cheapest fitting setup, Qwen/Qwen3-235B-A22B uses about 118 GB of VRAM per GPU for FP8 weights and about 6.3 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run Qwen/Qwen3-235B-A22B with INT4 quantization?
Yes, Qwen/Qwen3-235B-A22B has INT4 weight figures. Its weights take about 130 GB in INT4 versus about 236 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →