Best local LLMs by hardware tier
Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.
Looking for a coding model? See the best local LLM for coding by VRAM.
NVIDIA 16 GB
Midrange NVIDIA cards. Handles 13B comfortably and 30B-class with tight quants. Ranked from 65 reports.
| # | Model family | Best variant tested | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | Qwen3.8Alibaba 27B · 27B · NVFP4 · on V100 16GB | 27B · 27B · NVFP4 on V100 16GB | 48 | 219.1 |
| 2 | Qwen3.6Alibaba 35B-A3B · 35B-A3B · IQ4_XS-4.19bpw · on RTX 4070 Ti Super | 35B-A3B · 35B-A3B · IQ4_XS-4.19bpw | 11 | 110.2 |
| 3 | DeepSeek V4 FlashDeepSeek 284B-A13B · 284B-A13B · on RTX 5060 Ti 16GB | 284B-A13B · 284B-A13B | 2 | 11.0 |
| 4 | Agents-A1InternScience Uncensored-MTP-APEX · 35B-A3B · APEX Compact · on RTX 5070 Ti | Uncensored-MTP-APEX · 35B-A3B · APEX Compact on RTX 5070 Ti | 1 | 97.0 |
| 5 | Gemma 4Google DeepMind E4B · 7.5B · Q4_K_M · on RTX 5060 Ti 16GB | E4B · 7.5B · Q4_K_M | 1 | 73.9 |
| 6 | MuseMeta Glimmer · 30B · Q4_K_XL · on RTX 5060 Ti 16GB | Glimmer · 30B · Q4_K_XL | 1 | 18.0 |
| 7 | Qwen3Alibaba 30B-A3B · 30B-A3B · float8 · on RTX 5060 Ti 16GB | 30B-A3B · 30B-A3B · float8 | 1 | 52.0 |
How we rank
A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.