Best local LLMs by hardware tier
Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.
Looking for a coding model? See the best local LLM for coding by VRAM.
NVIDIA 32 GB+ workstation
Workstation and datacenter cards. 70B-class in a single device. Ranked from 143 reports.
| # | Model family | Best variant tested | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | Qwen3.8Alibaba Flash-Next · 125B-A6B · NVFP4 · on RTX PRO 6000 Max-Q | Flash-Next · 125B-A6B · NVFP4 | 61 | 240.0 |
| 2 | DeepSeek V4 FlashDeepSeek 284B-A13B · 284B-A13B · W4A16-FP8 · on H100 80GB | 284B-A13B · 284B-A13B · W4A16-FP8 on H100 80GB | 32 | 193.0 |
| 3 | Qwen3.6Alibaba 35B-A3B · 35B-A3B · Q4 · on RTX 5090 | 35B-A3B · 35B-A3B · Q4 on RTX 5090 | 14 | 215.1 |
| 4 | Gemma 4Google DeepMind 26B-A4B · 26B-A4B · AWQ-4bit · on RTX 5090 | 26B-A4B · 26B-A4B · AWQ-4bit on RTX 5090 | 6 | 578.0 |
| 5 | MuseMeta Glimmer · 30B · UD-Q5_K_M · on RTX 5090 | Glimmer · 30B · UD-Q5_K_M on RTX 5090 | 7 | 253.0 |
| 6 | Ling-3.0Ant Group INT4 · on DGX Spark | INT4 on DGX Spark | 4 | 40.9 |
| 7 | GLM-5.2Zhipu AI NVFP4 · on DGX Spark | NVFP4 on DGX Spark | 3 | 24.0 |
| 8 | Qwen3.5Alibaba 9B · 9B · on RTX 5090 | 9B · 9B on RTX 5090 | 2 | — |
| 9 | Qwen3Alibaba 8B · 8B · BF16 · on RTX 5090 | 8B · 8B · BF16 on RTX 5090 | 1 | 159.0 |
| 10 | DeepSeek V4 ProDeepSeek — · on RTX PRO 6000 Max-Q | — | 2 | 7.1 |
| 11 | Tencent-HY3Tencent 295B-A21B · 295B-A21B · Q6_K · on RTX Pro 6000 Blackwell | 295B-A21B · 295B-A21B · Q6_K | 1 | 57.4 |
| 12 | GLM-5.3Zhipu AI Flash · UD-Q4_K_XL · on RTX 5090 | Flash · UD-Q4_K_XL on RTX 5090 | 1 | 24.5 |
| 13 | Aurora1.0 0.1B · 0.1B · on RTX Pro 6000 Blackwell | 0.1B · 0.1B | 1 | — |
| 14 | HobbyLM 0.5B · 0.5B · on H200 | 0.5B · 0.5B on H200 | 1 | — |
| 15 | Kimi K3Moonshot AI — · on DGX Spark | — on DGX Spark | 1 | 20.0 |
| 16 | Mimo 2.5Xiaomi — · on RTX Pro 6000 Blackwell | — | 1 | — |
| 17 | Qwen2.5Alibaba 7B · 7B · on RTX 5090 | 7B · 7B on RTX 5090 | 1 | — |
| 18 | Llama 3.2Meta 1B · 1B · on RTX Pro 6000 Blackwell | 1B · 1B | 1 | — |
| 19 | Kimi K2.5Moonshot AI IQ3_M · on RTX 5090 | IQ3_M on RTX 5090 | 1 | — |
| 20 | DeepSeek V3DeepSeek — · on DGX Spark | — on DGX Spark | 1 | — |
| 21 | GLM-5.1Zhipu AI NVFP4 · on DGX Spark | NVFP4 on DGX Spark | 1 | — |
How we rank
A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.