gemma-4-31B-it-FP8-block (RedHatAI) is a 31.3B parameters open-weight model tracked on AI Hashrate. At 8K context we estimate roughly 24.48 GB VRAM for Q4 (weights 17.22 GB + KV 6.26 GB + 1 GB overhead) and 69.86 GB for FP16. About 24 GPUs/accelerators in our catalog fully fit this model at Q4 under that context assumption. Fastest Q4 configs: B200 192GB SXM ≈ 162.6 tok/s (estimated); MI300X 192GB ≈ 107.8 tok/s (estimated); H200 141GB SXM5 ≈ 97.6 tok/s (estimated); H100 80GB SXM5 ≈ 68.1 tok/s (estimated); Gaudi 3 128GB ≈ 66.6 tok/s (estimated). Use the table below for VRAM fit and tok/s per dollar (list MSRP). Estimates follow memory-bandwidth math; measured rows override when present. See methodology for details. Methodology.