Best local LLMs by hardware tier
Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.
Looking for a coding model? See the best local LLM for coding by VRAM.
NVIDIA 8–12 GB
Entry-level NVIDIA discrete cards. Comfortable home for 7B–13B models in Q4. Ranked from 41 reports.
| # | Model family | Best variant tested | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | Qwen3.8Alibaba 27B · 27B · Q4_K_XL · on RTX 3060 12GB | 27B · 27B · Q4_K_XL | 21 | 50.6 |
| 2 | Qwen3.6Alibaba 35B-A3B · 35B-A3B · on CMP 170HX | 35B-A3B · 35B-A3B on CMP 170HX | 9 | 210.0 |
| 3 | DeepSeek V4 FlashDeepSeek 284B-A13B · 284B-A13B · Q4_K_XL · on CMP 170HX | 284B-A13B · 284B-A13B · Q4_K_XL on CMP 170HX | 3 | 29.0 |
| 4 | Gemma 4Google DeepMind E2B Instruct · 5.1B · Q4_K_M · on RTX 3060 12GB | E2B Instruct · 5.1B · Q4_K_M | 3 | 60.0 |
| 5 | LFM2.5Liquid AI 1.2B-Instruct · 1.2B · on RTX 4050 6GB | 1.2B-Instruct · 1.2B on RTX 4050 6GB | 1 | 129.0 |
| 6 | Bonsai-27B 27B · 27B · 1-bit · on RTX 3060 Laptop 6GB | 27B · 27B · 1-bit | 1 | 30.0 |
| 7 | Qwen3.5Alibaba 13B · 13B · on RTX 3060 12GB | 13B · 13B | 1 | 12.0 |
| 8 | DeepSeek V4DeepSeek Flash · 284B · W8A8 · on RTX 2080 Ti | Flash · 284B · W8A8 on RTX 2080 Ti | 1 | — |
| 9 | KAT-Coder V2.5-DevKwaipilot — · on RTX 3060 12GB | — | 1 | 14.2 |
How we rank
A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.