llamaperf

Best local LLMs by hardware tier

Rankings only make sense once you fix the hardware. Pick a tier below — the leaderboard re-sorts to the models the community actually runs there, weighted by report count, fastest observed tokens-per-second, and recency.

Looking for a coding model? See the best local LLM for coding by VRAM.

Hardware tier
Model size

NVIDIA 32 GB+ workstation

Workstation and datacenter cards. 70B-class in a single device. Ranked from 143 reports.

#Model familyReportsFastest t/s
1Qwen3.8Alibaba
Flash-Next · 125B-A6B · NVFP4 · on RTX PRO 6000 Max-Q
61240.0
2DeepSeek V4 FlashDeepSeek
284B-A13B · 284B-A13B · W4A16-FP8 · on H100 80GB
32193.0
3Qwen3.6Alibaba
35B-A3B · 35B-A3B · Q4 · on RTX 5090
14215.1
4Gemma 4Google DeepMind
26B-A4B · 26B-A4B · AWQ-4bit · on RTX 5090
6578.0
5MuseMeta
Glimmer · 30B · UD-Q5_K_M · on RTX 5090
7253.0
6Ling-3.0Ant Group
INT4 · on DGX Spark
440.9
7GLM-5.2Zhipu AI
NVFP4 · on DGX Spark
324.0
8Qwen3.5Alibaba
9B · 9B · on RTX 5090
2
9Qwen3Alibaba
8B · 8B · BF16 · on RTX 5090
1159.0
10DeepSeek V4 ProDeepSeek
· on RTX PRO 6000 Max-Q
27.1
11Tencent-HY3Tencent
295B-A21B · 295B-A21B · Q6_K · on RTX Pro 6000 Blackwell
157.4
12GLM-5.3Zhipu AI
Flash · UD-Q4_K_XL · on RTX 5090
124.5
13Aurora1.0
0.1B · 0.1B · on RTX Pro 6000 Blackwell
1
14HobbyLM
0.5B · 0.5B · on H200
1
15Kimi K3Moonshot AI
· on DGX Spark
120.0
16Mimo 2.5Xiaomi
· on RTX Pro 6000 Blackwell
1
17Qwen2.5Alibaba
7B · 7B · on RTX 5090
1
18Llama 3.2Meta
1B · 1B · on RTX Pro 6000 Blackwell
1
19Kimi K2.5Moonshot AI
IQ3_M · on RTX 5090
1
20DeepSeek V3DeepSeek
· on DGX Spark
1
21GLM-5.1Zhipu AI
NVFP4 · on DGX Spark
1

How we rank

A single global "best models" list doesn't really exist — what runs well on a 5090 is often unrunnable on a 4060, and a 7B that screams on an M3 Max is usually a poor pick on an H100. So we fix the hardware first, then rank the families that actually have community reports on it. The score blends popularity (log-scaled report count), fastest observed tokens-per-second normalized within the bucket, recency (90-day half-life), and a small bias for rows where we know the variant + quant + GPU cleanly. Click into a family for the full breakdown of records.