AI Hashrate

AI Hashrate · RTX 5090 32GB

Which LLM should you run on the RTX 5090 32GB?

32 GB · 1792 GB/s · 575 W

Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.

The short answer

gemma-4-31B-it — about 54.8 tok/s decode

Also fits

ModelQ4 VRAMEst. tok/s
gemma-4-26B-A4B-it15.4 GB443.1
Qwen3.6-35B-A3B21.3 GB561.3
Qwen3.6-27B17.1 GB60.6
Qwen3.5-35B-A3B20.6 GB561.3
gpt-oss-20b13.2 GB467.7
Qwen3.5-9B6.3 GB188.1
Qwen3.5-27B16.6 GB62.4

Run it

llama-server \
  -hf <GGUF_REPO>:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --host 127.0.0.1 --port 8080

These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.

Get an exact answer for your machine →