Best local LLMs by hardware tier
Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.
Looking for a coding model? See the best local LLM for coding by VRAM.
Apple Silicon 48–96 GB
Mid-range Apple Silicon. 70B-class fits in Q4 with bandwidth-limited speed. Ranked from 32 reports.
| # | Model family | Best variant tested | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | Qwen3.8Alibaba T5 · on M3 Max 48GB | T5 on M3 Max 48GB | 11 | 38.0 |
| 2 | DeepSeek V4 FlashDeepSeek 9B distill · 9B · Q4_K_M · on M5 Pro 48GB | 9B distill · 9B · Q4_K_M on M5 Pro 48GB | 8 | 44.0 |
| 3 | Qwen3.6Alibaba 27B · 27B · 4-bit · on M5 Max 64GB | 27B · 27B · 4-bit on M5 Max 64GB | 6 | 63.0 |
| 4 | MuseMeta Glimmer · 30B · 8-bit · on M4 Pro 48GB | Glimmer · 30B · 8-bit on M4 Pro 48GB | 2 | 18.0 |
| 5 | Gemma 4Google DeepMind 26B-A4B · 26B-A4B · on M5 Max 64GB | 26B-A4B · 26B-A4B on M5 Max 64GB | 1 | 97.0 |
| 6 | Qwen3Alibaba 8B · 8B · Q4 · on M2 Max 96GB | 8B · 8B · Q4 on M2 Max 96GB | 1 | 43.0 |
| 7 | LingBot-World-V2 causal-fast · 1.3B · 8-bit · on M4 Pro 48GB | causal-fast · 1.3B · 8-bit on M4 Pro 48GB | 1 | — |
| 8 | GLM-5.3Zhipu AI JANG · on M4 Pro 48GB | JANG on M4 Pro 48GB | 1 | 1.8 |
| 9 | DeepSeek V4.1 Flash — · on M2 Max 64GB | — on M2 Max 64GB | 1 | 2.0 |
How we rank
A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.