Best local LLMs by hardware tier
Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.
Looking for a coding model? See the best local LLM for coding by VRAM.
NVIDIA 24 GB consumer
The community sweet spot — 30B-class in Q4 with headroom for context. Ranked from 110 reports.
| # | Model family | Best variant tested | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | Qwen3.8Alibaba 27B · 27B · on RTX 3090 | 27B · 27B on RTX 3090 | 62 | 382.0 |
| 2 | DeepSeek V4 FlashDeepSeek 284B-A13B · 284B-A13B · on RTX 4090 | 284B-A13B · 284B-A13B on RTX 4090 | 17 | 180.0 |
| 3 | Qwen3.6Alibaba 35B-A3B · 35B-A3B · ninfer quant · on RTX 3090 | 35B-A3B · 35B-A3B · ninfer quant on RTX 3090 | 16 | 170.7 |
| 4 | Gemma 4Google DeepMind 26B-A4B · 26B-A4B · Q4_K_M · on RTX 4090 | 26B-A4B · 26B-A4B · Q4_K_M on RTX 4090 | 2 | 149.6 |
| 5 | GLM-5.2Zhipu AI 744B-A40B · 744B-A40B · UD-IQ2_M · on RTX 3090 | 744B-A40B · 744B-A40B · UD-IQ2_M on RTX 3090 | 2 | 7.3 |
| 6 | Qwen2.5Alibaba 27B · 27B · Q6_K_XL · on RTX 3090 | 27B · 27B · Q6_K_XL on RTX 3090 | 2 | 70.0 |
| 7 | GLM-5.3Zhipu AI — · on RTX 3090 | — on RTX 3090 | 2 | 55.0 |
| 8 | MuseMeta Glimmer · 30B · Q4_K_XL · on RTX 3090 | Glimmer · 30B · Q4_K_XL on RTX 3090 | 1 | 94.0 |
| 9 | Qwen3Alibaba 235B-A22B · 235B-A22B · Q4_K_M · on RTX 3090 | 235B-A22B · 235B-A22B · Q4_K_M on RTX 3090 | 1 | 7.5 |
| 10 | GLM-4.5-Air 106B-A12B · 106B-A12B · on RTX 3090 | 106B-A12B · 106B-A12B on RTX 3090 | 1 | — |
| 11 | DeepSeek R1DeepSeek Q3 · on RTX 3090 | Q3 on RTX 3090 | 1 | 2.0 |
| 12 | DeepSeek V4DeepSeek — · on RTX 3090 | — on RTX 3090 | 1 | 15.0 |
| 13 | Chandra — · on L4 | — on L4 | 1 | — |
| 14 | DeepSeek V3DeepSeek Q3 · on RTX 3090 | Q3 on RTX 3090 | 1 | — |
How we rank
A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.