Qwen/Qwen3-Coder-Next VRAM requirements

80B params, Hybrid attention (full + sliding-window layers)

Qwen3-Coder-Next needs about 80 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 1× NVIDIA B200. It supports 64 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200164
NVIDIA H200 SXM164
NVIDIA H100 NVL19
NVIDIA RTX PRO 6000 Blackwell Server Edition114
NVIDIA H20114
NVIDIA H100 SXM264
NVIDIA A100 80GB SXM264
NVIDIA L40S464
32k context
GPUGPUsConcurrent users
NVIDIA B200164
NVIDIA H200 SXM141
NVIDIA H100 NVL13
NVIDIA RTX PRO 6000 Blackwell Server Edition15
NVIDIA H2015
NVIDIA H100 SXM264
NVIDIA A100 80GB SXM256
NVIDIA L40S444
128k context
GPUGPUsConcurrent users
NVIDIA B200124
NVIDIA H200 SXM111
NVIDIA H100 NVL10
NVIDIA RTX PRO 6000 Blackwell Server Edition11
NVIDIA H2011
NVIDIA H100 SXM219
NVIDIA A100 80GB SXM216
NVIDIA L40S412

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
BF16about 159 GB1x NVIDIA B200
FP8about 80 GB1x NVIDIA B200
INT4 (AWQ)about 44 GB1x NVIDIA B200

Model notes

Attention
Hybrid attention, full + sliding-window layers
Architecture
Mixture-of-Experts
Layers
48
Hidden size
2,048
Vocabulary
151,936
Native precision
bfloat16

At 32k context, 26% of this card's usable memory is working as conversation cache.

Frequently asked questions

Will Qwen/Qwen3-Coder-Next run on a single H100?
No, a single H100 SXM cannot hold Qwen/Qwen3-Coder-Next at 32k context with FP8 weights. The model requires multiple GPUs.
What is the cheapest GPU setup for Qwen/Qwen3-Coder-Next?
The cheapest fitting setup at 32k context is 1x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does Qwen/Qwen3-Coder-Next need at 128k context?
At 128k context on the cheapest fitting setup, Qwen/Qwen3-Coder-Next uses about 80 GB of VRAM per GPU for FP8 weights and about 1.7 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run Qwen/Qwen3-Coder-Next with INT4 quantization?
Yes, Qwen/Qwen3-Coder-Next has INT4 weight figures. Its weights take about 44 GB in INT4 versus about 80 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →