Qwen3.6-35B-A3B-FP8 (Qwen) is a 36.0B parameters MoE (36.0B active per token) open-weight model tracked on AI Hashrate. At 8K context we estimate roughly 28.0 GB VRAM for Q4 (weights 19.8 GB + KV 7.2 GB + 1 GB overhead) and 80.2 GB for FP16. About 24 GPUs/accelerators in our catalog fully fit this model at Q4 under that context assumption. Fastest Q4 configs: B200 192GB SXM ≈ 141.4 tok/s (estimated); MI300X 192GB ≈ 93.7 tok/s (estimated); H200 141GB SXM5 ≈ 84.8 tok/s (estimated); H100 80GB SXM5 ≈ 59.2 tok/s (estimated); Gaudi 3 128GB ≈ 57.9 tok/s (estimated). Use the table below for VRAM fit and tok/s per dollar (list MSRP). Estimates follow memory-bandwidth math; measured rows override when present. See methodology for details. Methodology.