2026 Local LLM GPU Buying Guide: What Actually Runs What

July 2026 · Data from aihashrate.stream — 73 models, 63 GPUs, 100+ measured anchors

Short version: For smooth Q4 7–9B chat, get 12GB+. For 30B+ dense models, 24GB+ with offload. For MoE models, VRAM budget for total params, not just active params. The RTX 4090 is still the best consumer card you can just buy on Amazon right now. If you're on a budget, a used RTX 3060 12GB punches way above its weight.

This is not a lab leaderboard. These numbers come from memory-bandwidth estimation calibrated against 100+ community-measured tok/s anchors. They're directionally right — but your specific backend, quant engine, and batch size will move the needle. Use this to narrow your options, then check the interactive tool with your exact model and context length.


1. VRAM Classes: What Actually Fits (Q4, 8K Context)

Every few days someone on r/LocalLLaMA asks: "Can I run X on Y VRAM?" Here's the real answer. These numbers use the standard formula: weights + KV cache + 1GB overhead. Q4 quantization, 8K context. For longer context, add more KV budget. For FP16, roughly double the weight budget.

If a model lands in "borderline", it means the weights fit but the KV cache + overhead pushes total over — you might need to drop context or offload layers. Models are ranked by HuggingFace download popularity.

VRAMExample GPUsModels that fit (Q4/8K)Borderline
48GB RTX 6000 Ada, L40S, RTX A6000 45/73 models — everything up to Qwen3.6-35B-A3B MoE, gemma-4-31B dense FP8 Llama-3.1-70B, Qwen2.5-72B KV pushes over
32GB RTX 5090, MacBook Pro M1 Max 32GB, Arc Pro B70 43/73 models Same as 48GB — 70B+ dense still borderline
24GB RTX 4090, RTX 3090, RX 7900 XTX 37/73 models — covers all 8B, most MoE A3B variants, gemma-4-26B Qwen3.6-35B-A3B-FP8, Qwen2.5-32B tight at 8K ctx
16GB RTX 5080, RTX 5070 Ti, RTX 4080 Super, RX 9070 XT 24/73 models — all 8B, gemma-4-26B MoE, gpt-oss-20b Qwen3.6-27B, gemma-4-26B-A4B AWQ barely fits
12GB RTX 4070, RTX 3060, Arc B580 18/73 models — all 7–9B, gemma-4-E2B, Qwen3-8B Qwen3-14B, gpt-oss-20b needs offload
8GB RTX 4060, RTX 5060, RX 7600, Tesla P4 16/73 models — 7–9B Q4 fits, but it's tight gemma-4-12B, OLMo-2-13B won't fit at 8K

Want to check a specific combo? Pick GPU + model on ai-hashrate and it'll show exact VRAM breakdown, tok/s estimate, and whether it's measured or estimated.

2. Measured Tok/s Anchors: Real Numbers, Not Marketing

These are real measurements from community testers, not estimates. Grouped by GPU. Each GPU runs different models, so the range reflects both model size and quant variation. Higher = better, but "usable for chat" starts around 20 tok/s. Smooth chat needs 30+.

GPUAnchorsAvg tok/sMax tok/sMin tok/sVerdict
RTX 5090 32GB616528220Excellent — fastest consumer card measured
RTX 5080 16GB3168205136Excellent — but 16GB VRAM limits model choice
RTX 5070 Ti 16GB3161190132Great
RTX 6000 Ada 48GB2153167139Great — but $6,800
RTX 5070 12GB2110120100Good
RTX 4090 24GB11922253Great — most anchors, best value 24GB card
RX 9070 XT 16GB39410389Good — AMD catching up
RTX 3090 24GB58116228Solid — used market king
RTX 4080 Super 16GB310318658Good
RTX 4070 Ti Super 16GB311619954Good
RTX 3060 12GB56012323Usable — budget champion, ~60 tok/s for 7B models
RX 7900 XTX 24GB4566830Usable — good VRAM, slower than NVIDIA at same tier
Mac mini M4 32GB2314121Adequate — unified memory works, just slower
RTX 4060 Ti 16GB2385026Slow — VRAM is nice, bandwidth is the bottleneck
Key takeaway
Bandwidth is everything. The RTX 4090 (1,008 GB/s) and RTX 5090 (1,792 GB/s) dominate because they have the memory bandwidth to push tokens fast. The Mac mini M4 has 32GB but only 120 GB/s — it fits bigger models, but they run slower. VRAM determines what fits. Bandwidth determines how fast. Both matter.

3. The MoE VRAM Trap — Everyone Gets This Wrong

Mixture-of-Experts models advertise small "active parameters" — Qwen3.6-35B-A3B activates only 3B per token. Sounds like it should fit on a potato. Doesn't work that way.

The rule: VRAM must hold ALL weights, not just active experts. Speed scales with active params. But if you don't have VRAM for all 35B weights, the model won't load — period.

MoE ModelActive ParamsTotal ParamsRatioVRAM needed (Q4)Fits on 24GB?
gemma-4-26B-A4B-it3.8B25.2B0.15~14.2 GBYes
Qwen3.6-35B-A3B3.0B36.0B0.08~19.7 GBYes (tight)
Qwen3.6-35B-A3B-FP836.0B36.0B1.00~36.3 GBNo
gpt-oss-20b3.6B21.5B0.17~12.3 GBYes

Bottom line: When someone says "this 35B MoE only activates 3B, so it's basically a 3B model" — they're wrong about VRAM. It activates like a 3B model (fast), but it consumes VRAM like a 35B model. The website at ai-hashrate handles this automatically — MoE models show both active and total params in the estimate.

4. Tokens Per Dollar: Where's the Value?

If you just want the most tok/s for your money — regardless of VRAM — here's the ranking. Uses Q4/8K estimate for a typical 8B model, divided by MSRP.

RankGPUVRAMEst. tok/sMSRPtok/s per $1KNotes
1Tesla P4 8GB815$80186Used server pull, no display output
2Arc B580 12GB1235$249142Intel's budget king, 12GB for $249
3Arc B570 10GB1030$219135
4Arc A770 16GB1644$34912516GB for $349 is hard to beat
5RTX 5060 8GB835$299116But 8GB VRAM limits you
6RX 7800 XT 16GB1648$49997Good AMD value pick
7RTX 5070 12GB1252$54995
8RTX 5070 Ti 16GB1670$74993
9RTX 3060 12GB1228$32985Used market ≈ $200 makes it #1 value
10RTX 4090 24GB24126$1,59979Lower per-dollar, but fits 2× more models

Value rankings use MSRP. Used market prices shift everything — an RTX 3090 at $700 used beats almost everything new. But we can't track used prices in real time, so we use MSRP as the anchor.

5. Quick Picks: What to Buy at Each Budget

BudgetBest PickWhy
< $300 Arc B580 12GB or used RTX 3060 12GB 12GB runs all 7–9B models at Q4 smoothly. 3060 has CUDA (broader LLM backend support). B580 has faster bandwidth but Intel drivers are still maturing.
$300–500 RTX 4060 Ti 16GB or Arc A770 16GB 16GB VRAM opens up gemma-4-26B MoE and gpt-oss-20b. 4060 Ti has CUDA but slower bandwidth. A770 is cheaper but Intel quirks.
$500–800 RTX 5070 12GB or used RTX 3090 24GB RTX 5070 is fast but only 12GB. Used 3090 gives you 24GB — fits 37/73 models — at similar price. This is the sweet spot if you can find a good used 3090.
$800–1,200 RTX 5070 Ti 16GB or RX 9070 XT 16GB Fastest 16GB cards. Both handle all 8B models at blazing speed. NVIDIA has better software ecosystem.
$1,200–2,000 RTX 4090 24GB Still the king. 1,008 GB/s bandwidth. Fits 37/73 models. 100+ tok/s on 8B models. Amazon buy link works. If you can find one at close to MSRP, buy it.
$2,000+ RTX 5090 32GB Fastest measured consumer GPU (avg 165 tok/s). 32GB fits 43/73 models. But at $1,999 MSRP (and likely more on the street), it's a luxury pick.
The used RTX 3090 play
If you find a used RTX 3090 24GB for $600–800, it's the best deal in local LLM hardware right now. 24GB VRAM, 936 GB/s bandwidth, CUDA ecosystem, and 5 measured anchors averaging 81 tok/s on Q4 models. The only downside: 350W TDP, so make sure your PSU can handle it, and it runs hot — budget for good case airflow.

6. What About Apple Silicon?

Apple's unified memory is a different beast. An M4 Mac mini with 32GB has only 120 GB/s bandwidth — but it can run models that need 20+ GB of VRAM because the memory is shared. At Q4, it fits everything up to gemma-4-31B and Qwen3.6-35B-A3B MoE.

The trade-off: speed. You're getting 20–40 tok/s instead of 100+. For interactive chat, 30 tok/s is fine. For batch processing or speculative decoding, NVIDIA wins by a mile.

Measured anchors: Mac mini M4 32GB does 31 tok/s avg on Q4 models. M2 Ultra gives you more bandwidth (800 GB/s) but costs $5,000+ — at which point you could buy two RTX 4090s.


Where These Numbers Come From

Every estimate on this page uses the same formula as our methodology:

These are ballpark numbers, not lab benchmarks. Your actual tok/s depends on your backend (llama.cpp vs vLLM vs MLX), quant engine, batch size, and system load. The relative rankings are more reliable than the absolute numbers.

Try the interactive tool →   Pick any GPU + model, switch quant and context, see real-time estimates.


MSRP is manufacturer list price. Amazon prices are whatever you see after clicking. Some links may earn us a commission at no extra cost to you (Amazon Associates). Missing a GPU or model? Send it in. Got real tok/s numbers from your setup? Share them and we'll add them as measured anchors.