llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX Pro 4000 Blackwell
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
71.4 tokens/s generation
quant:
AWQ (AWQ)
kv:
fp8_e5m2

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User benchmarks Qwen2.5-32B-Instruct-AWQ at 71.37 tok/s median decode on 2x RTX PRO 4000 Blackwell. Setup is SGLang 0.5.19 with AWQ weights, NGRAM speculative decoding (K=5), tensor parallelism 2 over PCIe Gen4, and FP8 KV cache. NGRAM speculative decoding gives a 1.28x median speedup over the 55.60 tok/s baseline, with high inter-prompt variance (std dev 19.18 tok/s). Per-domain decode ranges are 65-107 tok/s for JSON, 60-226 tok/s for code, and 58-82 tok/s for prose. FP8 KV cache alone measured 54.77 tok/s, a 1.5% regression.

Sep 26, 2026
Tone: positive
reported speed:
67.0 tokens/s generation · 785.0 tokens/s prompt processing
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 67 t/s with MTP3 speculative decoding enabled, against 24.4 tok/s without MTP. Setup uses an INT8 KV cache with group-64. The figure comes from a 128K NIAH benchmark with 130,048 prompt tokens, where MTP acceptance was 100% on a deterministic answer. Roughly 727 MiB of VRAM was left.

Sep 7, 2026
Tone: positive
reported speed:
33.9 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8:27b at 33.91 t/s on 2x RTX PRO 4000s and an RTX PRO 2000, up from 12.2 t/s. Splitting the GPUs across VMs produced the gain. User also reports muse-glimmer:30B at 22.61 t/s on the same setup, up from 14.3 t/s.

Sep 7, 2026
Showing 1–3 of 3
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23