Best GPUs for 30B local LLMs
30B-class models in Q4 fit comfortably in 24GB of VRAM with room for a useful context window — the sweet spot for a single consumer GPU. Apple Silicon Macs with 32GB+ unified memory also handle them well. Ranked from community reports.
Ranked from 486 community reports on llamaperf.
Ranked by community reports
| # | GPU | VRAM | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | RTX 3090nvidia | 24GB | 87 | 382.0 |
| 2 | RTX 5090nvidia | 32GB | 67 | 578.0 |
| 3 | AMD Strix Halo 128GBamd | 128GB | 35 | 153.3 |
| 4 | RTX Pro 6000 Blackwellnvidia | 96GB | 27 | 177.0 |
| 5 | RTX 5060 Ti 16GBnvidia | 16GB | 26 | 76.7 |
| 6 | DGX Sparknvidia | 128GB | 25 | 181.0 |
| 7 | M5 Max 128GBapple | 128GB | 18 | 133.6 |
| 8 | Radeon AI PRO R9700 32GBamd | 32GB | 17 | 280.0 |
| 9 | RTX 4090nvidia | 24GB | 17 | 180.0 |
| 10 | RX 7900 XTXamd | 24GB | 15 | 100.0 |
| 11 | RTX 5070 Tinvidia | 16GB | 14 | 115.0 |
| 12 | RTX 5080nvidia | 16GB | 7 | 75.0 |
| 13 | RTX PRO 6000 Max-Qnvidia | 96GB | 6 | 240.0 |
| 14 | V100 32GBnvidia | 32GB | 6 | 218.0 |
| 15 | RTX 3080 20GBnvidia | 20GB | 6 | 57.5 |
| 16 | M2 Max 96GBapple | 96GB | 6 | 43.0 |
| 17 | RTX 4060 Ti 16GBnvidia | 16GB | 6 | 32.5 |
| 18 | AMD MI50 32GBamd | 32GB | 6 | 15.5 |
| 19 | H100 80GBnvidia | 80GB | 5 | 193.0 |
| 20 | RTX 6000nvidia | 48GB | 5 | 150.0 |
| 21 | RTX 4070 Ti Supernvidia | 16GB | 5 | 110.2 |
| 22 | M4 Max 128GBapple | 128GB | 5 | 72.5 |
| 23 | M4 Pro 48GBapple | 48GB | 5 | 20.3 |
| 24 | M5 Max 64GBapple | 64GB | 4 | 97.0 |
| 25 | RX 9070amd | 16GB | 4 | 73.0 |
| 26 | M1 Max 64GBapple | 64GB | 4 | 21.0 |
| 27 | M3 Ultra 512GBapple | 512GB | 4 | 20.0 |
| 28 | M5 Pro 64GBapple | 64GB | 4 | 20.0 |
| 29 | RTX 4080nvidia | 16GB | 3 | 56.5 |
| 30 | M2 Ultra 192GBapple | 192GB | 3 | 28.0 |
| 31 | H200nvidia | 141GB | 3 | 4.8 |
| 32 | V100 16GBnvidia | 16GB | 2 | 219.1 |
| 33 | M5 Pro 48GBapple | 48GB | 2 | 44.0 |
| 34 | M3 Ultra 256GBapple | 256GB | 2 | 37.4 |
| 35 | M4 Max 64GBapple | 64GB | 2 | 36.0 |
| 36 | M1 Ultra 128GBapple | 128GB | 2 | 31.2 |
| 37 | M4 32GBapple | 32GB | 2 | 22.0 |
| 38 | RTX A6000 48GBnvidia | 48GB | 2 | 17.2 |
| 39 | M1 Max 32GBapple | 32GB | 2 | 15.8 |
| 40 | M3 Max 96GBapple | 96GB | 2 | 12.7 |
| 41 | M3 Max 128GBapple | 128GB | 2 | 5.5 |
| 42 | M5 32GBapple | 32GB | 2 | 1.0 |
| 43 | RTX 3090 Tinvidia | 24GB | 1 | 100.0 |
| 44 | RTX 4080 Supernvidia | 16GB | 1 | 59.0 |
| 45 | A100 80GBnvidia | 80GB | 1 | 56.8 |
| 46 | RX 7900 GRE 16GBamd | 16GB | 1 | 51.9 |
| 47 | RTX Pro 4500 Blackwell 32GBnvidia | 32GB | 1 | 45.2 |
| 48 | RX 6800 16GBamd | 16GB | 1 | 45.0 |
| 49 | M3 Ultra 192GBapple | 192GB | 1 | 43.0 |
| 50 | M3 Max 48GBapple | 48GB | 1 | 38.0 |
| 51 | M4 16GBapple | 16GB | 1 | 20.0 |
| 52 | M3 Pro 36GBapple | 36GB | 1 | 17.7 |
| 53 | T4 16GBnvidia | 16GB | 1 | 17.6 |
| 54 | A100 40GBnvidia | 40GB | 1 | 16.1 |
| 55 | M5 16GBapple | 16GB | 1 | 9.0 |
| 56 | M2 Pro 32GBapple | 32GB | 1 | 8.6 |
| 57 | AMD Threadripper 256GBamd | 256GB | 1 | 7.5 |
| 58 | M2 Max 64GBapple | 64GB | 1 | 2.0 |
| 59 | Instinct MI300X 192GBamd | 192GB | 1 | — |
| 60 | M3 Ultra 96GBapple | 96GB | 1 | — |
| 61 | L4nvidia | 24GB | 1 | — |
Models that fit
No reports yet
These match the profile but nobody has submitted a report yet.
What to look for
24GB cards are the sweet spot
RTX 3090s and 4090s (both 24GB) hold a 30B-class model in Q4 with plenty of headroom for an 8–16K context. This is arguably the best price/capability point in local LLM inference today — you get most of the quality of a 70B model at a fraction of the hardware cost.
16GB cards work with tighter quants
An RTX 4060 Ti 16GB or RTX 4070 Ti Super 16GB can run 30B models at Q3/Q4 with shorter contexts, though you'll feel the squeeze with longer prompts. Q3 quants noticeably hurt quality on most models — Q4 is the practical floor.
Frequently asked
What's the best GPU for a 30B local LLM?
RTX 3090 (used) or RTX 4090 (new) — both 24GB — are the standard recommendations. They hold a 30B model in Q4 with headroom for a useful context window and run at 25–50 tokens-per-second on most engines.
Can a 16GB GPU run 30B models?
Yes, with caveats. Q3/Q4 quants of 30B-class models fit in ~14–17GB depending on the architecture. You'll have less context room and may need to lower precision further than ideal. A 24GB card is meaningfully better.
How we rank
Hardware is sorted by the number of community submissions on llamaperf — a proxy for how widely each card is used in practice for local LLM inference. Within that, we surface the fastest tokens-per-second observed on each as a quality signal. Submissions come primarily from r/LocalLLaMA discussions and direct user uploads. Nothing here is sponsored or affiliate-driven.