AI Hashrate

AI Hashrate · RTX 3090 24GB

Which LLM should you run on the RTX 3090 24GB?

24 GB · 936 GB/s · 350 W

Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.

The short answer

gemma-4-26B-A4B-it — about 210.4 tok/s decode

Also fits

ModelQ4 VRAMEst. tok/s
gpt-oss-20b13.2 GB222.1
Qwen3.5-9B6.3 GB89.3
Qwen3.6-35B-A3B21.3 GB266.5
Ornith-1.0-9B6.3 GB88.8
Qwen3-8B6.7 GB97.5
gemma-4-12b-it8.2 GB66.9
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1618.2 GB266.5

Run it

llama-server \
  -hf <GGUF_REPO>:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --host 127.0.0.1 --port 8080

These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.

Get an exact answer for your machine →