685B params, Multi-head latent attention (MLA)
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 8 | 64 |
| NVIDIA H200 SXM | 8 | 64 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 8 | 2 |
| NVIDIA H20 | 8 | 2 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 64 |
| NVIDIA H100 NVL | 16 (verify server configuration) | 64 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 64 |
| NVIDIA L40S | no fit |
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 8 | 64 |
| NVIDIA H200 SXM | 8 | 28 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 46 |
| NVIDIA H100 NVL | 16 (verify server configuration) | 64 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 38 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 16 (verify server configuration) | 64 |
| NVIDIA H20 | 16 (verify server configuration) | 64 |
| NVIDIA L40S | no fit |
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 8 | 16 |
| NVIDIA H200 SXM | 8 | 7 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 11 |
| NVIDIA H100 NVL | 16 (verify server configuration) | 16 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 9 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 16 (verify server configuration) | 17 |
| NVIDIA H20 | 16 (verify server configuration) | 17 |
| NVIDIA L40S | no fit |
FP8 weights, SGLang defaults, estimates.
| Precision | Weights | Smallest fitting setup at 32k |
|---|---|---|
| FP8 | about 673 GB | 8x NVIDIA B200 |
| INT4 (AWQ) | about 370 GB | 4x NVIDIA B200 |
MLA keeps one compact shared KV copy per GPU, so adding GPUs does not shrink the cache, but the copy itself is small for a model this size.
Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.
Open in calculatorDon't want to run this yourself? We deploy and operate it for you. Book a call →