llamaperf

Best GPUs for 30B local LLMs

30B-class models in Q4 fit comfortably in 24GB of VRAM with room for a useful context window — the sweet spot for a single consumer GPU. Apple Silicon Macs with 32GB+ unified memory also handle them well. Ranked from community reports.

Ranked from 486 community reports on llamaperf.

Ranked by community reports

#GPUVRAMReportsFastest t/s
1RTX 3090nvidia24GB87382.0
2RTX 5090nvidia32GB67578.0
3AMD Strix Halo 128GBamd128GB35153.3
4RTX Pro 6000 Blackwellnvidia96GB27177.0
5RTX 5060 Ti 16GBnvidia16GB2676.7
6DGX Sparknvidia128GB25181.0
7M5 Max 128GBapple128GB18133.6
8Radeon AI PRO R9700 32GBamd32GB17280.0
9RTX 4090nvidia24GB17180.0
10RX 7900 XTXamd24GB15100.0
11RTX 5070 Tinvidia16GB14115.0
12RTX 5080nvidia16GB775.0
13RTX PRO 6000 Max-Qnvidia96GB6240.0
14V100 32GBnvidia32GB6218.0
15RTX 3080 20GBnvidia20GB657.5
16M2 Max 96GBapple96GB643.0
17RTX 4060 Ti 16GBnvidia16GB632.5
18AMD MI50 32GBamd32GB615.5
19H100 80GBnvidia80GB5193.0
20RTX 6000nvidia48GB5150.0
21RTX 4070 Ti Supernvidia16GB5110.2
22M4 Max 128GBapple128GB572.5
23M4 Pro 48GBapple48GB520.3
24M5 Max 64GBapple64GB497.0
25RX 9070amd16GB473.0
26M1 Max 64GBapple64GB421.0
27M3 Ultra 512GBapple512GB420.0
28M5 Pro 64GBapple64GB420.0
29RTX 4080nvidia16GB356.5
30M2 Ultra 192GBapple192GB328.0
31H200nvidia141GB34.8
32V100 16GBnvidia16GB2219.1
33M5 Pro 48GBapple48GB244.0
34M3 Ultra 256GBapple256GB237.4
35M4 Max 64GBapple64GB236.0
36M1 Ultra 128GBapple128GB231.2
37M4 32GBapple32GB222.0
38RTX A6000 48GBnvidia48GB217.2
39M1 Max 32GBapple32GB215.8
40M3 Max 96GBapple96GB212.7
41M3 Max 128GBapple128GB25.5
42M5 32GBapple32GB21.0
43RTX 3090 Tinvidia24GB1100.0
44RTX 4080 Supernvidia16GB159.0
45A100 80GBnvidia80GB156.8
46RX 7900 GRE 16GBamd16GB151.9
47RTX Pro 4500 Blackwell 32GBnvidia32GB145.2
48RX 6800 16GBamd16GB145.0
49M3 Ultra 192GBapple192GB143.0
50M3 Max 48GBapple48GB138.0
51M4 16GBapple16GB120.0
52M3 Pro 36GBapple36GB117.7
53T4 16GBnvidia16GB117.6
54A100 40GBnvidia40GB116.1
55M5 16GBapple16GB19.0
56M2 Pro 32GBapple32GB18.6
57AMD Threadripper 256GBamd256GB17.5
58M2 Max 64GBapple64GB12.0
59Instinct MI300X 192GBamd192GB1
60M3 Ultra 96GBapple96GB1
61L4nvidia24GB1

Models that fit

No reports yet

These match the profile but nobody has submitted a report yet.

What to look for

24GB cards are the sweet spot

RTX 3090s and 4090s (both 24GB) hold a 30B-class model in Q4 with plenty of headroom for an 8–16K context. This is arguably the best price/capability point in local LLM inference today — you get most of the quality of a 70B model at a fraction of the hardware cost.

16GB cards work with tighter quants

An RTX 4060 Ti 16GB or RTX 4070 Ti Super 16GB can run 30B models at Q3/Q4 with shorter contexts, though you'll feel the squeeze with longer prompts. Q3 quants noticeably hurt quality on most models — Q4 is the practical floor.

Frequently asked

What's the best GPU for a 30B local LLM?

RTX 3090 (used) or RTX 4090 (new) — both 24GB — are the standard recommendations. They hold a 30B model in Q4 with headroom for a useful context window and run at 25–50 tokens-per-second on most engines.

Can a 16GB GPU run 30B models?

Yes, with caveats. Q3/Q4 quants of 30B-class models fit in ~14–17GB depending on the architecture. You'll have less context room and may need to lower precision further than ideal. A 24GB card is meaningfully better.

How we rank

Hardware is sorted by the number of community submissions on llamaperf — a proxy for how widely each card is used in practice for local LLM inference. Within that, we surface the fastest tokens-per-second observed on each as a quality signal. Submissions come primarily from r/LocalLLaMA discussions and direct user uploads. Nothing here is sponsored or affiliate-driven.

See also