← AI Hashrate 中文

AI21-Jamba-Large-1.5

ai21labs · 398.6B · ctx NoneK · other · HuggingFace

AI21-Jamba-Large-1.5 (ai21labs) is a 398.6B parameters open-weight model tracked on AI Hashrate. At 8K context we estimate roughly 299.95 GB VRAM for Q4 (weights 219.23 GB + KV 79.72 GB + 1 GB overhead) and 877.92 GB for FP16. About 1 GPUs/accelerators in our catalog fully fit this model at Q4 under that context assumption. Fastest Q4 configs: DGX Spark (GH200 480GB) ≈ 0.8 tok/s (estimated). Use the table below for VRAM fit and tok/s per dollar (list MSRP). Estimates follow memory-bandwidth math; measured rows override when present. See methodology for details. Methodology.

13 hardware configs — sorted by tok/s. Default fit context 8K.

HardwareVRAMQuanttok/sFits?tok/s/$
B200 192GB SXM192Q414.9 est.No0.0
MI300X 192GB192Q49.9 est.No0.001
H200 141GB SXM5141Q44.8 est.No0.0
Gaudi 3 128GB128Q42.7 est.No0.0
Mac Studio M2 Ultra 192GB192Q41.5 est.No0.0
H100 80GB SXM580Q41.1 est.No0.0
DGX Spark (GH200 480GB)480Q40.8 est.Yes0.0
A100 80GB SXM480Q40.7 est.No0.0
MacBook Pro M4 Max 128GB128Q40.5 est.No0.0
Mac Studio M3 Ultra 96GB96Q40.4 est.No0.0
MacBook Pro M3 Max 128GB128Q40.3 est.No0.0
DGX Spark (GH200 480GB)480FP160.2 est.No0.0
Ryzen AI Max+ 395 128GB96Q40.1 est.No0.0

Fit checks

12 popular retail GPUs; every row in the table above also links its Fits? verdict to the full fit check.