llamaperf

Best local LLMs by hardware tier

Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.

Looking for a coding model? See the best local LLM for coding by VRAM.

Hardware tier
Model size

Apple 128 GB+ & Strix Halo

Large unified-memory rigs. Big MoE models, very long contexts. Ranked from 73 reports.

#Model familyReportsFastest t/s
1Qwen3.8Alibaba
27B · 27B · Q4_K_M · on AMD Strix Halo 128GB
27153.3
2Qwen3.6Alibaba
35B-A3B · 35B-A3B · Q8 XL · on M5 Max 128GB
1579.4
3DeepSeek V4 FlashDeepSeek
284B-A13B · 284B-A13B · MXFP4 · on M3 Ultra 192GB
1043.0
4Nex-N2.5-mini
MLX-4bit · on M5 Max 128GB
2133.6
5GLM-5.3Zhipu AI
Flash · Q4 · on M3 Ultra 256GB
337.4
6Qwen3.5Alibaba
Opus-Reasoning · 122B-A10B · Q4_K_XL · on AMD Strix Halo 128GB
249.0
7Gemma 4Google DeepMind
31B · 31B · on M5 Max 128GB
47.5
8LFM2.5Liquid AI
2.6B · 2.6B · Q4_K_M · on AMD Strix Halo 128GB
1113.0
9Tencent-HY3Tencent
295B-A21B · 295B-A21B · UD128 · on M5 Max 128GB
132.4
10MuseMeta
Glimmer · 30B · UD-Q2_K_XL · on M4 Max 128GB
125.0
11Qwen3-Coder-Next
UD-Q6_K_XL · on AMD Strix Halo 128GB
136.8
12GLM-4.7Zhipu AI
· on AMD Strix Halo 128GB
140.0
13StepFun 3.7StepFun
Q4_K_S · on M5 Max 128GB
133.9
14DeepSeek V4.1 Flash
oQ4e · on M3 Ultra 512GB
119.7
15Minimax M3MiniMax
4bit · on M3 Ultra 256GB
117.3
16DeepSeek V4DeepSeek
IQ3_XXS · on M1 Ultra 128GB
116.0
17DeepSeek V3DeepSeek
· on M2 Ultra 192GB
128.0

How we rank

A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.