1599B params, Compressed recurrent state
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 8 | 64 |
| NVIDIA H200 SXM | 8 | 23 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 36 |
| NVIDIA H100 NVL | 16 (verify server configuration) | 64 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 24 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 16 (verify server configuration) | 64 |
| NVIDIA H20 | 16 (verify server configuration) | 64 |
| NVIDIA L40S | no fit |
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 8 | 32 |
| NVIDIA H200 SXM | 8 | 6 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 9 |
| NVIDIA H100 NVL | 16 (verify server configuration) | 17 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 6 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 16 (verify server configuration) | 18 |
| NVIDIA H20 | 16 (verify server configuration) | 18 |
| NVIDIA L40S | no fit |
| GPU | GPUs | Concurrent users |
|---|---|---|
| NVIDIA B200 | 8 | 8 |
| NVIDIA H200 SXM | 8 | 1 |
| NVIDIA H100 SXM | 16 (verify server configuration) | 2 |
| NVIDIA H100 NVL | 16 (verify server configuration) | 4 |
| NVIDIA A100 80GB SXM | 16 (verify server configuration) | 1 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 16 (verify server configuration) | 4 |
| NVIDIA H20 | 16 (verify server configuration) | 4 |
| NVIDIA L40S | no fit |
FP8 weights, SGLang defaults, estimates.
| Precision | Weights | Smallest fitting setup at 32k |
|---|---|---|
| FP8 | about 851 GB | 8x NVIDIA B200 |
| INT4 (AWQ) | about 851 GB | 8x NVIDIA B200 |
This model keeps a fixed-size recurrent state per conversation instead of a growing token cache, so very long chats cost little extra memory.
Want a different context length, GPU, or quantization? The calculator runs the same numbers live, preloaded with this model.
Open in calculatorDon't want to run this yourself? We deploy and operate it for you. Book a call →