llamaperf

Best local LLMs by hardware tier

Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.

Looking for a coding model? See the best local LLM for coding by VRAM.

Hardware tier
Model size

Apple Silicon 24–36 GB

Lower-tier Apple Silicon. 13B–30B with MLX or llama.cpp Metal. Ranked from 8 reports.

#Model familyReportsFastest t/s
1Qwen3.8Alibaba
Flash-Next · 125B-A6B · Q4 · on M4 32GB
622.0
2DeepSeek V4 FlashDeepSeek
284B-A13B · 284B-A13B · 4bit · on M5 32GB
11.0
3Gemma 4Google DeepMind
· on M5 32GB
1

How we rank

A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.