AI Hashrate

AI Hashrate · MacBook Pro M4 Max 128GB

Which LLM should you run on the MacBook Pro M4 Max 128GB?

128 GB · 546 GB/s · 160 W

Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.

The short answer

Laguna S 2.1 — about 96.2 tok/s decode

Also fits

ModelQ4 VRAMEst. tok/s
gpt-oss-120b68.2 GB148.3
Qwen3.6-35B-A3B20.8 GB252.1
gemma-4-26B-A4B-it14.9 GB199
Qwen3-Next-80B-A3B-Instruct46.8 GB252.1
Qwen3.5-122B-A10B69 GB75.6
Qwen3.5-35B-A3B20.1 GB252.1
Llama-4-Scout-17B-16E-Instruct64.9 GB44.5

Run it

llama-server \
  -hf <GGUF_REPO>:Q4_K_M \
  -c 8192 \
  -ngl 99 \
  --host 127.0.0.1 --port 8080

These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.

Get an exact answer for your machine →