gpt-oss-120b needs about 65 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 1× NVIDIA B200. It supports 64 concurrent conversations with SGLang defaults.
GPU requirements by context length
8k context
GPU
GPUs
Concurrent users
NVIDIA B200
1
64
NVIDIA H200 SXM
1
64
NVIDIA H100 SXM
1
36
NVIDIA H100 NVL
1
64
NVIDIA A100 80GB SXM
1
5
NVIDIA RTX PRO 6000 Blackwell Server Edition
1
64
NVIDIA H20
1
64
NVIDIA L40S
2
64
32k context
GPU
GPUs
Concurrent users
NVIDIA B200
1
64
NVIDIA H200 SXM
1
64
NVIDIA H100 SXM
1
9
NVIDIA H100 NVL
1
30
NVIDIA A100 80GB SXM
1
1
NVIDIA RTX PRO 6000 Blackwell Server Edition
1
33
NVIDIA H20
1
33
NVIDIA L40S
2
21
128k context
GPU
GPUs
Concurrent users
NVIDIA B200
1
39
NVIDIA H200 SXM
1
21
NVIDIA H100 SXM
1
2
NVIDIA H100 NVL
1
7
NVIDIA RTX PRO 6000 Blackwell Server Edition
1
8
NVIDIA H20
1
8
NVIDIA A100 80GB SXM
2
27
NVIDIA L40S
2
5
FP8 weights, SGLang defaults, estimates.
Weights by precision
Weight memory
Precision
Weights
Smallest fitting setup at 32k
FP8
about 65 GB
1x NVIDIA B200
INT4 (AWQ)
about 65 GB
1x NVIDIA B200
Model notes
Attention
Hybrid attention, full + sliding-window layers
Architecture
Mixture-of-Experts
Sliding window
128 tokens on 18 of 36 layers
Layers
36
Hidden size
2,880
Vocabulary
201,088
The sliding-window design caps this model's memory appetite, past 128 tokens, 18 of its 36 layers stop charging for longer conversations.
Frequently asked questions
Will openai/gpt-oss-120b run on a single H100?
Yes, openai/gpt-oss-120b fits on a single H100 SXM at 32k context with FP8 weights, supporting 9 concurrent conversations.
What is the cheapest GPU setup for openai/gpt-oss-120b?
The cheapest fitting setup at 32k context is 1x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does openai/gpt-oss-120b need at 128k context?
At 128k context on the cheapest fitting setup, openai/gpt-oss-120b uses about 65 GB of VRAM per GPU for FP8 weights and about 2.4 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run openai/gpt-oss-120b with INT4 quantization?
Yes, openai/gpt-oss-120b has INT4 weight figures. Its weights take about 65 GB in INT4 versus about 65 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.
Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.