2026 Local LLM GPU Buying Guide: What Actually Runs What
Short version: For smooth Q4 7–9B chat, get 12GB+. For 30B+ dense models, 24GB+ with offload. For MoE models, VRAM budget for total params, not just active params. The RTX 4090 is still the best consumer card you can just buy on Amazon right now. If you're on a budget, a used RTX 3060 12GB punches way above its weight.
This is not a lab leaderboard. These numbers come from memory-bandwidth estimation calibrated against 100+ community-measured tok/s anchors. They're directionally right — but your specific backend, quant engine, and batch size will move the needle. Use this to narrow your options, then check the interactive tool with your exact model and context length.
1. VRAM Classes: What Actually Fits (Q4, 8K Context)
Every few days someone on r/LocalLLaMA asks: "Can I run X on Y VRAM?" Here's the real answer. These numbers use the standard formula: weights + KV cache + 1GB overhead. Q4 quantization, 8K context. For longer context, add more KV budget. For FP16, roughly double the weight budget.
If a model lands in "borderline", it means the weights fit but the KV cache + overhead pushes total over — you might need to drop context or offload layers. Models are ranked by HuggingFace download popularity.
| VRAM | Example GPUs | Models that fit (Q4/8K) | Borderline |
|---|---|---|---|
| 48GB | RTX 6000 Ada, L40S, RTX A6000 | 45/73 models — everything up to Qwen3.6-35B-A3B MoE, gemma-4-31B dense FP8 | Llama-3.1-70B, Qwen2.5-72B KV pushes over |
| 32GB | RTX 5090, MacBook Pro M1 Max 32GB, Arc Pro B70 | 43/73 models | Same as 48GB — 70B+ dense still borderline |
| 24GB | RTX 4090, RTX 3090, RX 7900 XTX | 37/73 models — covers all 8B, most MoE A3B variants, gemma-4-26B | Qwen3.6-35B-A3B-FP8, Qwen2.5-32B tight at 8K ctx |
| 16GB | RTX 5080, RTX 5070 Ti, RTX 4080 Super, RX 9070 XT | 24/73 models — all 8B, gemma-4-26B MoE, gpt-oss-20b | Qwen3.6-27B, gemma-4-26B-A4B AWQ barely fits |
| 12GB | RTX 4070, RTX 3060, Arc B580 | 18/73 models — all 7–9B, gemma-4-E2B, Qwen3-8B | Qwen3-14B, gpt-oss-20b needs offload |
| 8GB | RTX 4060, RTX 5060, RX 7600, Tesla P4 | 16/73 models — 7–9B Q4 fits, but it's tight | gemma-4-12B, OLMo-2-13B won't fit at 8K |
Want to check a specific combo? Pick GPU + model on ai-hashrate and it'll show exact VRAM breakdown, tok/s estimate, and whether it's measured or estimated.
2. Measured Tok/s Anchors: Real Numbers, Not Marketing
These are real measurements from community testers, not estimates. Grouped by GPU. Each GPU runs different models, so the range reflects both model size and quant variation. Higher = better, but "usable for chat" starts around 20 tok/s. Smooth chat needs 30+.
| GPU | Anchors | Avg tok/s | Max tok/s | Min tok/s | Verdict |
|---|---|---|---|---|---|
| RTX 5090 32GB | 6 | 165 | 282 | 20 | Excellent — fastest consumer card measured |
| RTX 5080 16GB | 3 | 168 | 205 | 136 | Excellent — but 16GB VRAM limits model choice |
| RTX 5070 Ti 16GB | 3 | 161 | 190 | 132 | Great |
| RTX 6000 Ada 48GB | 2 | 153 | 167 | 139 | Great — but $6,800 |
| RTX 5070 12GB | 2 | 110 | 120 | 100 | Good |
| RTX 4090 24GB | 11 | 92 | 225 | 3 | Great — most anchors, best value 24GB card |
| RX 9070 XT 16GB | 3 | 94 | 103 | 89 | Good — AMD catching up |
| RTX 3090 24GB | 5 | 81 | 162 | 28 | Solid — used market king |
| RTX 4080 Super 16GB | 3 | 103 | 186 | 58 | Good |
| RTX 4070 Ti Super 16GB | 3 | 116 | 199 | 54 | Good |
| RTX 3060 12GB | 5 | 60 | 123 | 23 | Usable — budget champion, ~60 tok/s for 7B models |
| RX 7900 XTX 24GB | 4 | 56 | 68 | 30 | Usable — good VRAM, slower than NVIDIA at same tier |
| Mac mini M4 32GB | 2 | 31 | 41 | 21 | Adequate — unified memory works, just slower |
| RTX 4060 Ti 16GB | 2 | 38 | 50 | 26 | Slow — VRAM is nice, bandwidth is the bottleneck |
3. The MoE VRAM Trap — Everyone Gets This Wrong
Mixture-of-Experts models advertise small "active parameters" — Qwen3.6-35B-A3B activates only 3B per token. Sounds like it should fit on a potato. Doesn't work that way.
The rule: VRAM must hold ALL weights, not just active experts. Speed scales with active params. But if you don't have VRAM for all 35B weights, the model won't load — period.
| MoE Model | Active Params | Total Params | Ratio | VRAM needed (Q4) | Fits on 24GB? |
|---|---|---|---|---|---|
| gemma-4-26B-A4B-it | 3.8B | 25.2B | 0.15 | ~14.2 GB | Yes |
| Qwen3.6-35B-A3B | 3.0B | 36.0B | 0.08 | ~19.7 GB | Yes (tight) |
| Qwen3.6-35B-A3B-FP8 | 36.0B | 36.0B | 1.00 | ~36.3 GB | No |
| gpt-oss-20b | 3.6B | 21.5B | 0.17 | ~12.3 GB | Yes |
Bottom line: When someone says "this 35B MoE only activates 3B, so it's basically a 3B model" — they're wrong about VRAM. It activates like a 3B model (fast), but it consumes VRAM like a 35B model. The website at ai-hashrate handles this automatically — MoE models show both active and total params in the estimate.
4. Tokens Per Dollar: Where's the Value?
If you just want the most tok/s for your money — regardless of VRAM — here's the ranking. Uses Q4/8K estimate for a typical 8B model, divided by MSRP.
| Rank | GPU | VRAM | Est. tok/s | MSRP | tok/s per $1K | Notes |
|---|---|---|---|---|---|---|
| 1 | Tesla P4 8GB | 8 | 15 | $80 | 186 | Used server pull, no display output |
| 2 | Arc B580 12GB | 12 | 35 | $249 | 142 | Intel's budget king, 12GB for $249 |
| 3 | Arc B570 10GB | 10 | 30 | $219 | 135 | |
| 4 | Arc A770 16GB | 16 | 44 | $349 | 125 | 16GB for $349 is hard to beat |
| 5 | RTX 5060 8GB | 8 | 35 | $299 | 116 | But 8GB VRAM limits you |
| 6 | RX 7800 XT 16GB | 16 | 48 | $499 | 97 | Good AMD value pick |
| 7 | RTX 5070 12GB | 12 | 52 | $549 | 95 | |
| 8 | RTX 5070 Ti 16GB | 16 | 70 | $749 | 93 | |
| 9 | RTX 3060 12GB | 12 | 28 | $329 | 85 | Used market ≈ $200 makes it #1 value |
| 10 | RTX 4090 24GB | 24 | 126 | $1,599 | 79 | Lower per-dollar, but fits 2× more models |
Value rankings use MSRP. Used market prices shift everything — an RTX 3090 at $700 used beats almost everything new. But we can't track used prices in real time, so we use MSRP as the anchor.
5. Quick Picks: What to Buy at Each Budget
| Budget | Best Pick | Why |
|---|---|---|
| < $300 | Arc B580 12GB or used RTX 3060 12GB | 12GB runs all 7–9B models at Q4 smoothly. 3060 has CUDA (broader LLM backend support). B580 has faster bandwidth but Intel drivers are still maturing. |
| $300–500 | RTX 4060 Ti 16GB or Arc A770 16GB | 16GB VRAM opens up gemma-4-26B MoE and gpt-oss-20b. 4060 Ti has CUDA but slower bandwidth. A770 is cheaper but Intel quirks. |
| $500–800 | RTX 5070 12GB or used RTX 3090 24GB | RTX 5070 is fast but only 12GB. Used 3090 gives you 24GB — fits 37/73 models — at similar price. This is the sweet spot if you can find a good used 3090. |
| $800–1,200 | RTX 5070 Ti 16GB or RX 9070 XT 16GB | Fastest 16GB cards. Both handle all 8B models at blazing speed. NVIDIA has better software ecosystem. |
| $1,200–2,000 | RTX 4090 24GB | Still the king. 1,008 GB/s bandwidth. Fits 37/73 models. 100+ tok/s on 8B models. Amazon buy link works. If you can find one at close to MSRP, buy it. |
| $2,000+ | RTX 5090 32GB | Fastest measured consumer GPU (avg 165 tok/s). 32GB fits 43/73 models. But at $1,999 MSRP (and likely more on the street), it's a luxury pick. |
6. What About Apple Silicon?
Apple's unified memory is a different beast. An M4 Mac mini with 32GB has only 120 GB/s bandwidth — but it can run models that need 20+ GB of VRAM because the memory is shared. At Q4, it fits everything up to gemma-4-31B and Qwen3.6-35B-A3B MoE.
The trade-off: speed. You're getting 20–40 tok/s instead of 100+. For interactive chat, 30 tok/s is fine. For batch processing or speculative decoding, NVIDIA wins by a mile.
Measured anchors: Mac mini M4 32GB does 31 tok/s avg on Q4 models. M2 Ultra gives you more bandwidth (800 GB/s) but costs $5,000+ — at which point you could buy two RTX 4090s.
Where These Numbers Come From
Every estimate on this page uses the same formula as our methodology:
- VRAM = weights + KV cache + 1GB overhead. Weights are params × bytes-per-param (Q4 ≈ 0.55 bytes, FP16 ≈ 2 bytes). KV cache scales with context length.
- Tok/s ≈ bandwidth / bytes-per-token × utilization factor. We use 0.35 for fitting models (conservative), lower for offloaded models. The formula is simple physics — memory bandwidth is the bottleneck, not compute.
- Measured anchors calibrate the estimates. When community testers report real tok/s numbers, we record them and use estimates for the rest. Look for "measured" vs "estimated" tags on the site.
These are ballpark numbers, not lab benchmarks. Your actual tok/s depends on your backend (llama.cpp vs vLLM vs MLX), quant engine, batch size, and system load. The relative rankings are more reliable than the absolute numbers.
Try the interactive tool → Pick any GPU + model, switch quant and context, see real-time estimates.
MSRP is manufacturer list price. Amazon prices are whatever you see after clicking. Some links may earn us a commission at no extra cost to you (Amazon Associates). Missing a GPU or model? Send it in. Got real tok/s numbers from your setup? Share them and we'll add them as measured anchors.