llamaperf

Best local LLMs by hardware tier

Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.

Looking for a coding model? See the best local LLM for coding by VRAM.

Hardware tier
Model size

NVIDIA 8–12 GB

Entry-level NVIDIA discrete cards. Comfortable home for 7B–13B models in Q4. Ranked from 41 reports.

#Model familyReportsFastest t/s
1Qwen3.8Alibaba
27B · 27B · Q4_K_XL · on RTX 3060 12GB
2150.6
2Qwen3.6Alibaba
35B-A3B · 35B-A3B · on CMP 170HX
9210.0
3DeepSeek V4 FlashDeepSeek
284B-A13B · 284B-A13B · Q4_K_XL · on CMP 170HX
329.0
4Gemma 4Google DeepMind
E2B Instruct · 5.1B · Q4_K_M · on RTX 3060 12GB
360.0
5LFM2.5Liquid AI
1.2B-Instruct · 1.2B · on RTX 4050 6GB
1129.0
6Bonsai-27B
27B · 27B · 1-bit · on RTX 3060 Laptop 6GB
130.0
7Qwen3.5Alibaba
13B · 13B · on RTX 3060 12GB
112.0
8DeepSeek V4DeepSeek
Flash · 284B · W8A8 · on RTX 2080 Ti
1
9KAT-Coder V2.5-DevKwaipilot
· on RTX 3060 12GB
114.2

How we rank

A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.