Best local LLMs by hardware tier
Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.
Looking for a coding model? See the best local LLM for coding by VRAM.
Apple 128 GB+ & Strix Halo
Large unified-memory rigs. Big MoE models, very long contexts. Ranked from 73 reports.
| # | Model family | Best variant tested | Reports | Fastest t/s |
|---|---|---|---|---|
| 1 | Qwen3.8Alibaba 27B · 27B · Q4_K_M · on AMD Strix Halo 128GB | 27B · 27B · Q4_K_M | 27 | 153.3 |
| 2 | Qwen3.6Alibaba 35B-A3B · 35B-A3B · Q8 XL · on M5 Max 128GB | 35B-A3B · 35B-A3B · Q8 XL on M5 Max 128GB | 15 | 79.4 |
| 3 | DeepSeek V4 FlashDeepSeek 284B-A13B · 284B-A13B · MXFP4 · on M3 Ultra 192GB | 284B-A13B · 284B-A13B · MXFP4 | 10 | 43.0 |
| 4 | Nex-N2.5-mini MLX-4bit · on M5 Max 128GB | MLX-4bit on M5 Max 128GB | 2 | 133.6 |
| 5 | GLM-5.3Zhipu AI Flash · Q4 · on M3 Ultra 256GB | Flash · Q4 | 3 | 37.4 |
| 6 | Qwen3.5Alibaba Opus-Reasoning · 122B-A10B · Q4_K_XL · on AMD Strix Halo 128GB | Opus-Reasoning · 122B-A10B · Q4_K_XL | 2 | 49.0 |
| 7 | Gemma 4Google DeepMind 31B · 31B · on M5 Max 128GB | 31B · 31B on M5 Max 128GB | 4 | 7.5 |
| 8 | LFM2.5Liquid AI 2.6B · 2.6B · Q4_K_M · on AMD Strix Halo 128GB | 2.6B · 2.6B · Q4_K_M | 1 | 113.0 |
| 9 | Tencent-HY3Tencent 295B-A21B · 295B-A21B · UD128 · on M5 Max 128GB | 295B-A21B · 295B-A21B · UD128 on M5 Max 128GB | 1 | 32.4 |
| 10 | MuseMeta Glimmer · 30B · UD-Q2_K_XL · on M4 Max 128GB | Glimmer · 30B · UD-Q2_K_XL on M4 Max 128GB | 1 | 25.0 |
| 11 | Qwen3-Coder-Next UD-Q6_K_XL · on AMD Strix Halo 128GB | UD-Q6_K_XL | 1 | 36.8 |
| 12 | GLM-4.7Zhipu AI — · on AMD Strix Halo 128GB | — | 1 | 40.0 |
| 13 | StepFun 3.7StepFun Q4_K_S · on M5 Max 128GB | Q4_K_S on M5 Max 128GB | 1 | 33.9 |
| 14 | DeepSeek V4.1 Flash oQ4e · on M3 Ultra 512GB | oQ4e | 1 | 19.7 |
| 15 | Minimax M3MiniMax 4bit · on M3 Ultra 256GB | 4bit | 1 | 17.3 |
| 16 | DeepSeek V4DeepSeek IQ3_XXS · on M1 Ultra 128GB | IQ3_XXS | 1 | 16.0 |
| 17 | DeepSeek V3DeepSeek — · on M2 Ultra 192GB | — | 1 | 28.0 |
How we rank
A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.