1059B params, Multi-head latent attention (MLA)
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 4 | 39 |
| NVIDIA H200 SXM | 8 | 64 |
| NVIDIA H100 NVL | 8 | 29 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 8 | 36 |
| NVIDIA H20 | 8 | 36 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 64 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 64 |
| NVIDIA L40S | 16 (verify server configuration) | 2 |
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 4 | 9 |
| NVIDIA H200 SXM | 8 | 37 |
| NVIDIA H100 NVL | 8 | 7 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 8 | 9 |
| NVIDIA H20 | 8 | 9 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 54 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 46 |
| NVIDIA L40S | no fit |
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 4 | 2 |
| NVIDIA H200 SXM | 8 | 9 |
| NVIDIA H100 NVL | 8 | 1 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 8 | 2 |
| NVIDIA H20 | 8 | 2 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 13 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 11 |
| NVIDIA L40S | no fit |
FP8 weights, SGLang defaults, estimates.
| Precision | Weights | Smallest fitting setup at 32k |
|---|---|---|
| FP8 | about 595 GB | 4x NVIDIA B200 |
| INT4 (AWQ) | about 582 GB | 4x NVIDIA B200 |
MLA keeps one compact shared KV copy per GPU, so adding GPUs does not shrink the cache, but the copy itself is small for a model this size.
Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.
Open in calculatorDon't want to run this yourself? We deploy and operate it for you. Book a call →