llamaperf

Best local LLMs by hardware tier

Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.

Looking for a coding model? See the best local LLM for coding by VRAM.

Hardware tier
Model size

NVIDIA 24 GB consumer

The community sweet spot — 30B-class in Q4 with headroom for context. Ranked from 110 reports.

#Model familyReportsFastest t/s
1Qwen3.8Alibaba
27B · 27B · on RTX 3090
62382.0
2DeepSeek V4 FlashDeepSeek
284B-A13B · 284B-A13B · on RTX 4090
17180.0
3Qwen3.6Alibaba
35B-A3B · 35B-A3B · ninfer quant · on RTX 3090
16170.7
4Gemma 4Google DeepMind
26B-A4B · 26B-A4B · Q4_K_M · on RTX 4090
2149.6
5GLM-5.2Zhipu AI
744B-A40B · 744B-A40B · UD-IQ2_M · on RTX 3090
27.3
6Qwen2.5Alibaba
27B · 27B · Q6_K_XL · on RTX 3090
270.0
7GLM-5.3Zhipu AI
· on RTX 3090
255.0
8MuseMeta
Glimmer · 30B · Q4_K_XL · on RTX 3090
194.0
9Qwen3Alibaba
235B-A22B · 235B-A22B · Q4_K_M · on RTX 3090
17.5
10GLM-4.5-Air
106B-A12B · 106B-A12B · on RTX 3090
1
11DeepSeek R1DeepSeek
Q3 · on RTX 3090
12.0
12DeepSeek V4DeepSeek
· on RTX 3090
115.0
13Chandra
· on L4
1
14DeepSeek V3DeepSeek
Q3 · on RTX 3090
1

How we rank

A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.