Best GPUs for running local LLMs
Picking a GPU for local LLM inference comes down to VRAM (does the model fit?), memory bandwidth (how fast it generates), and software support. The list below is ranked by how many community reports each card has on llamaperf — a rough proxy for how heavily it gets used in practice — and surfaces the fastest tokens-per-second observed on each.
Ranked from 536 community reports on llamaperf.
Ranked by community reports
| # | GPU | VRAM | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | RTX 3090nvidia | 24GB | 87 | 382.0 |
| 2 | RTX 5090nvidia | 32GB | 67 | 578.0 |
| 3 | AMD Strix Halo 128GBamd | 128GB | 35 | 153.3 |
| 4 | RTX Pro 6000 Blackwellnvidia | 96GB | 27 | 177.0 |
| 5 | RTX 5060 Ti 16GBnvidia | 16GB | 26 | 76.7 |
| 6 | DGX Sparknvidia | 128GB | 25 | 181.0 |
| 7 | RTX 3060 12GBnvidia | 12GB | 24 | 70.0 |
| 8 | M5 Max 128GBapple | 128GB | 18 | 133.6 |
| 9 | Radeon AI PRO R9700 32GBamd | 32GB | 17 | 280.0 |
| 10 | RTX 4090nvidia | 24GB | 17 | 180.0 |
| 11 | RX 7900 XTXamd | 24GB | 15 | 100.0 |
| 12 | RTX 5070 Tinvidia | 16GB | 14 | 115.0 |
| 13 | RTX 5080nvidia | 16GB | 7 | 75.0 |
| 14 | RTX PRO 6000 Max-Qnvidia | 96GB | 6 | 240.0 |
| 15 | V100 32GBnvidia | 32GB | 6 | 218.0 |
| 16 | RTX 3080 20GBnvidia | 20GB | 6 | 57.5 |
| 17 | M2 Max 96GBapple | 96GB | 6 | 43.0 |
| 18 | RTX 4060 Ti 16GBnvidia | 16GB | 6 | 32.5 |
| 19 | AMD MI50 32GBamd | 32GB | 6 | 15.5 |
| 20 | CMP 170HXnvidia | 8GB | 5 | 210.0 |
| 21 | H100 80GBnvidia | 80GB | 5 | 193.0 |
| 22 | RTX 6000nvidia | 48GB | 5 | 150.0 |
| 23 | RTX 4070 Ti Supernvidia | 16GB | 5 | 110.2 |
| 24 | M4 Max 128GBapple | 128GB | 5 | 72.5 |
| 25 | M4 Pro 48GBapple | 48GB | 5 | 20.3 |
| 26 | M5 Max 64GBapple | 64GB | 4 | 97.0 |
| 27 | RX 9070amd | 16GB | 4 | 73.0 |
| 28 | RTX 4070nvidia | 12GB | 4 | 55.0 |
| 29 | M1 Max 64GBapple | 64GB | 4 | 21.0 |
| 30 | M3 Ultra 512GBapple | 512GB | 4 | 20.0 |
| 31 | M5 Pro 64GBapple | 64GB | 4 | 20.0 |
| 32 | Intel Arc Pro B70intel | 12GB | 3 | 70.5 |
| 33 | RTX 5070 Ti Laptop 12GBnvidia | 12GB | 3 | 59.0 |
| 34 | RTX 4080nvidia | 16GB | 3 | 56.5 |
| 35 | M2 Ultra 192GBapple | 192GB | 3 | 28.0 |
| 36 | H200nvidia | 141GB | 3 | 4.8 |
| 37 | V100 16GBnvidia | 16GB | 2 | 219.1 |
| 38 | RTX 2080 Tinvidia | 11GB | 2 | 45.0 |
| 39 | M5 Pro 48GBapple | 48GB | 2 | 44.0 |
| 40 | M3 Ultra 256GBapple | 256GB | 2 | 37.4 |
| 41 | M4 Max 64GBapple | 64GB | 2 | 36.0 |
| 42 | M1 Ultra 128GBapple | 128GB | 2 | 31.2 |
| 43 | M4 32GBapple | 32GB | 2 | 22.0 |
| 44 | RTX 5070nvidia | 12GB | 2 | 22.0 |
| 45 | RTX A6000 48GBnvidia | 48GB | 2 | 17.2 |
| 46 | M1 Max 32GBapple | 32GB | 2 | 15.8 |
| 47 | M3 Max 96GBapple | 96GB | 2 | 12.7 |
| 48 | M3 Max 128GBapple | 128GB | 2 | 5.5 |
| 49 | M5 32GBapple | 32GB | 2 | 1.0 |
| 50 | RTX 4050 6GBnvidia | 6GB | 1 | 129.0 |
| 51 | RTX 3090 Tinvidia | 24GB | 1 | 100.0 |
| 52 | RTX 4080 Supernvidia | 16GB | 1 | 59.0 |
| 53 | A100 80GBnvidia | 80GB | 1 | 56.8 |
| 54 | RX 7900 GRE 16GBamd | 16GB | 1 | 51.9 |
| 55 | RTX Pro 4500 Blackwell 32GBnvidia | 32GB | 1 | 45.2 |
| 56 | RX 6800 16GBamd | 16GB | 1 | 45.0 |
| 57 | M3 Ultra 192GBapple | 192GB | 1 | 43.0 |
| 58 | M3 Max 48GBapple | 48GB | 1 | 38.0 |
| 59 | RTX 3060 Laptop 6GBnvidia | 6GB | 1 | 30.0 |
| 60 | NVIDIA P102-100nvidia | 10GB | 1 | 23.5 |
| 61 | M4 16GBapple | 16GB | 1 | 20.0 |
| 62 | M3 Pro 36GBapple | 36GB | 1 | 17.7 |
| 63 | T4 16GBnvidia | 16GB | 1 | 17.6 |
| 64 | M1 8GBapple | 8GB | 1 | 17.5 |
| 65 | A100 40GBnvidia | 40GB | 1 | 16.1 |
| 66 | D700 12GBamd | 12GB | 1 | 11.0 |
| 67 | M5 16GBapple | 16GB | 1 | 9.0 |
| 68 | M2 Pro 32GBapple | 32GB | 1 | 8.6 |
| 69 | AMD Threadripper 256GBamd | 256GB | 1 | 7.5 |
| 70 | M2 8GBapple | 8GB | 1 | 5.5 |
| 71 | RX 5700 XT 8GBamd | 8GB | 1 | 2.7 |
| 72 | M2 Max 64GBapple | 64GB | 1 | 2.0 |
| 73 | Instinct MI300X 192GBamd | 192GB | 1 | — |
| 74 | M3 Ultra 96GBapple | 96GB | 1 | — |
| 75 | L4nvidia | 24GB | 1 | — |
No reports yet
These match the profile but nobody has submitted a report yet.
What to look for
VRAM is the gating constraint
Whether a model runs at all is decided by memory. A Q4_K_M quant of a 7B model needs ~5GB; a 13B needs ~8GB; a 30B needs ~20GB; a 70B needs ~40GB — plus headroom for context and KV cache. If the weights don't fit, generation either crawls (CPU offload) or fails outright.
Bandwidth determines tokens-per-second
Once weights fit, throughput is dominated by memory bandwidth, not raw FLOPs. An RTX 3090 (936 GB/s) and an RTX 4090 (1008 GB/s) are within ~10% of each other on inference-bound workloads despite the 4090's much larger compute budget. M-series Macs trade off here: massive memory pool, but Pro-tier bandwidth is closer to a midrange discrete card.
Software support gates which engines you can use
NVIDIA has CUDA kernels in every major engine (llama.cpp, vLLM, exllamav2, TensorRT-LLM). AMD support has improved sharply via ROCm but still trails on engine coverage. Apple Silicon is best-in-class for MLX and llama.cpp Metal but unsupported by vLLM. Match the engine you want to use to the hardware ecosystem.
Frequently asked
What is the best GPU for running local LLMs?
There is no single answer — it depends on which model size you want to run. For 7B–13B models, an RTX 3060 12GB or RTX 4060 Ti 16GB is enough. For 30B-class models, an RTX 3090 or 4090 (24GB) is the sweet spot. For 70B-class, you need 40GB+ of VRAM (RTX A6000, dual 3090s, or an M-series Mac with 64GB+ unified memory).
Is more VRAM or more compute better for local LLMs?
VRAM, by a wide margin. Inference throughput is memory-bandwidth bound, not compute bound. A card with enough VRAM to fit your model and decent bandwidth will outperform a faster GPU that has to offload weights to system memory.
Do I need an NVIDIA GPU for local LLMs?
No. AMD GPUs work via ROCm with most major engines, and Apple Silicon Macs run llama.cpp Metal and MLX natively. NVIDIA still has the broadest engine support and best out-of-the-box experience, but it's no longer the only option.
How is this list ranked?
By the number of community submissions on llamaperf for each GPU. More reports indicate a GPU is widely used in practice for local LLM inference. The fastest tokens-per-second observed on each is shown alongside as a quality signal.
How we rank
Hardware is sorted by the number of community submissions on llamaperf — a proxy for how widely each card is used in practice for local LLM inference. Within that, we surface the fastest tokens-per-second observed on each as a quality signal. Submissions come primarily from r/LocalLLaMA discussions and direct user uploads. Nothing here is sponsored or affiliate-driven.