AI Hashrate · MacBook Pro M4 Max 128GB
Which LLM should you run on the MacBook Pro M4 Max 128GB?
128 GB · 546 GB/s · 160 W
Every number below comes from the same engine behind the wizard: VRAM fit per quantisation, decode speed from memory bandwidth, and a command you can paste. Nothing is scraped or invented.
The short answer
Laguna S 2.1 — about 96.2 tok/s decode
Also fits
| Model | Q4 VRAM | Est. tok/s |
|---|---|---|
| gpt-oss-120b | 68.2 GB | 148.3 |
| Qwen3.6-35B-A3B | 20.8 GB | 252.1 |
| gemma-4-26B-A4B-it | 14.9 GB | 199 |
| Qwen3-Next-80B-A3B-Instruct | 46.8 GB | 252.1 |
| Qwen3.5-122B-A10B | 69 GB | 75.6 |
| Qwen3.5-35B-A3B | 20.1 GB | 252.1 |
| Llama-4-Scout-17B-16E-Instruct | 64.9 GB | 44.5 |
Run it
llama-server \
-hf <GGUF_REPO>:Q4_K_M \
-c 8192 \
-ngl 99 \
--host 127.0.0.1 --port 8080
These figures are modelled estimates from hardware specs and public reports, not benchmarks measured on this card. The wizard shows where every number comes from.