AI Hashrate · RTX 4090 24GB
Which LLM should you run on the RTX 4090 24GB?
24 GB · 1008 GB/s · 450 W
Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.
The short answer
gemma-4-26B-A4B-it — about 226.6 tok/s decode
Also fits
| Model | Q4 VRAM | Est. tok/s |
|---|---|---|
| gpt-oss-20b | 13.2 GB | 239.2 |
| Llama-3.1-8B-Lexi-Uncensored-V2 | 6.5 GB | 107.6 |
| Dolphin3.0-R1-Mistral-24B | 15.7 GB | 35.9 |
| Qwen3.5-9B | 6.3 GB | 96.2 |
| Cydonia-24B-v4.1 | 15.7 GB | 35.9 |
| Dolphin-Mistral-24B-Venice-Edition | 15.7 GB | 35.9 |
| Qwen3.8-27B-Uncensored | 16.9 GB | 31.3 |
Run it
llama-server \
-hf <GGUF_REPO>:Q4_K_M \
-c 8192 \
-ngl 99 \
--host 127.0.0.1 --port 8080
These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.