AI Hashrate

AI Hashrate · Ryzen AI Max+ 395 128GB

Which LLM should you run on the Ryzen AI Max+ 395 128GB?

96 GB · 215 GB/s · 120 W

Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.

The short answer

Qwen3.6-35B-A3B — about 66.5 tok/s decode

Also fits

ModelQ4 VRAMEst. tok/s
gemma-4-26B-A4B-it15.4 GB52.5
Qwen3-Next-80B-A3B-Instruct47.3 GB66.5
Qwen3.5-35B-A3B20.6 GB66.5
gpt-oss-20b13.2 GB55.4
Ornith-1.0-35B20.8 GB66.5
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1618.2 GB66.5
Qwen3-30B-A3B-Instruct-250718.8 GB66.5

Run it

llama-server \
  -hf <GGUF_REPO>:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --host 127.0.0.1 --port 8080

These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.

Get an exact answer for your machine →