google/gemma-4-31B-it VRAM requirements

33B params, Hybrid attention (full + sliding-window layers)

gemma-4-31B-it needs about 34 GB of GPU memory for FP8 weights. The smallest fitting setup at 32k context is 1× NVIDIA B200. It supports 64 concurrent conversations with SGLang defaults.

GPU requirements by context length

8k context
GPUGPUsConcurrent users
NVIDIA B200164
NVIDIA H200 SXM164
NVIDIA H100 SXM148
NVIDIA H100 NVL164
NVIDIA A100 80GB SXM142
NVIDIA L40S16
NVIDIA RTX PRO 6000 Blackwell Server Edition164
NVIDIA H20164
32k context
GPUGPUsConcurrent users
NVIDIA B200164
NVIDIA H200 SXM147
NVIDIA H100 SXM120
NVIDIA H100 NVL128
NVIDIA A100 80GB SXM118
NVIDIA L40S12
NVIDIA RTX PRO 6000 Blackwell Server Edition129
NVIDIA H20129
128k context
GPUGPUsConcurrent users
NVIDIA B200121
NVIDIA H200 SXM114
NVIDIA H100 SXM16
NVIDIA H100 NVL18
NVIDIA A100 80GB SXM15
NVIDIA RTX PRO 6000 Blackwell Server Edition18
NVIDIA H2018
NVIDIA L40S27

FP8 weights, SGLang defaults, estimates.

Weights by precision

Weight memory
PrecisionWeightsSmallest fitting setup at 32k
BF16about 63 GB1x NVIDIA B200
FP8about 34 GB1x NVIDIA B200
INT4 (AWQ)about 20 GB1x NVIDIA B200

Model notes

Attention
Hybrid attention, full + sliding-window layers
Architecture
Dense, 32.7B parameters
Sliding window
1,024 tokens on 50 of 60 layers
Layers
60
Hidden size
5,376
Vocabulary
262,144

The sliding-window design caps this model's memory appetite, past 1,024 tokens, 50 of its 60 layers stop charging for longer conversations.

Frequently asked questions

Will google/gemma-4-31B-it run on a single H100?
Yes, google/gemma-4-31B-it fits on a single H100 SXM at 32k context with FP8 weights, supporting 20 concurrent conversations.
What is the cheapest GPU setup for google/gemma-4-31B-it?
The cheapest fitting setup at 32k context is 1x NVIDIA B200 with FP8 weights and SGLang defaults.
How much VRAM does google/gemma-4-31B-it need at 128k context?
At 128k context on the cheapest fitting setup, google/gemma-4-31B-it uses about 34 GB of VRAM per GPU for FP8 weights and about 5.8 GB per conversation for the KV cache. The total VRAM needed depends on the GPU count and parallelism configuration.
Can I run google/gemma-4-31B-it with INT4 quantization?
Yes, google/gemma-4-31B-it has INT4 weight figures. Its weights take about 20 GB in INT4 versus about 34 GB in FP8. INT4 can make the model fit on fewer GPUs, but check inference quality for your workload.

Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.

Open in calculator

Don't want to run this yourself? We deploy and operate it for you. Book a call →