AI21-Jamba-Large-1.5 (ai21labs) is a 398.6B parameters open-weight model tracked on AI Hashrate. At 8K context we estimate roughly 299.95 GB VRAM for Q4 (weights 219.23 GB + KV 79.72 GB + 1 GB overhead) and 877.92 GB for FP16. About 1 GPUs/accelerators in our catalog fully fit this model at Q4 under that context assumption. Fastest Q4 configs: DGX Spark (GH200 480GB) ≈ 0.8 tok/s (estimated). Use the table below for VRAM fit and tok/s per dollar (list MSRP). Estimates follow memory-bandwidth math; measured rows override when present. See methodology for details. Methodology.
| Hardware | VRAM | Quant | tok/s | Fits? | tok/s/$ |
|---|---|---|---|---|---|
| B200 192GB SXM | 192 | Q4 | 14.9 est. | No | 0.0 |
| MI300X 192GB | 192 | Q4 | 9.9 est. | No | 0.001 |
| H200 141GB SXM5 | 141 | Q4 | 4.8 est. | No | 0.0 |
| Gaudi 3 128GB | 128 | Q4 | 2.7 est. | No | 0.0 |
| Mac Studio M2 Ultra 192GB | 192 | Q4 | 1.5 est. | No | 0.0 |
| H100 80GB SXM5 | 80 | Q4 | 1.1 est. | No | 0.0 |
| DGX Spark (GH200 480GB) | 480 | Q4 | 0.8 est. | Yes | 0.0 |
| A100 80GB SXM4 | 80 | Q4 | 0.7 est. | No | 0.0 |
| MacBook Pro M4 Max 128GB | 128 | Q4 | 0.5 est. | No | 0.0 |
| Mac Studio M3 Ultra 96GB | 96 | Q4 | 0.4 est. | No | 0.0 |
| MacBook Pro M3 Max 128GB | 128 | Q4 | 0.3 est. | No | 0.0 |
| DGX Spark (GH200 480GB) | 480 | FP16 | 0.2 est. | No | 0.0 |
| Ryzen AI Max+ 395 128GB | 96 | Q4 | 0.1 est. | No | 0.0 |