AI Hashrate

AI Hashrate · RTX 4090 24GB

Which LLM should you run on the RTX 4090 24GB?

24 GB · 1008 GB/s · 450 W

Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.

The short answer

gemma-4-26B-A4B-it — about 226.6 tok/s decode

Also fits

ModelQ4 VRAMEst. tok/s
gpt-oss-20b13.2 GB239.2
Qwen3.5-9B6.3 GB96.2
Ornith-1.0-9B6.3 GB95.7
Qwen3.6-27B17.1 GB31
Qwen3-8B6.7 GB105
gemma-4-12b-it8.2 GB72
Qwen3.6-35B-A3B21.3 GB287

Run it

llama-server \
  -hf <GGUF_REPO>:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --host 127.0.0.1 --port 8080

These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.

Get an exact answer for your machine →