llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX Pro 6000 Blackwell
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
200.9 tokens/s generation · 1943.0 tokens/s prompt processing
quant:
FP4 (FP8)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contexttool-usevision

User reports DeepSeek-V4.1-Flash at 200.9 t/s decode and 1,943.0 t/s prefill on 4× RTX PRO 6000 Blackwell 96GB GPUs. Setup is SGLang with FP4 routed experts and FP8 components, DSpark block 5 speculative decoding, 524,288 context, 64 GiB DDR5 cache, NVMe offload, 8 slots, GPUs capped at 275 W over PCIe without NVLink. The 200.9 t/s is the C1 single-stream figure from a 45-case sweep; total decode across 8 concurrent streams reached 713.5 t/s. The 4M populated-KV gate passed with 4,063,744 tokens across eight 500k inputs. Long-context answer quality and a one-hour soak remain unqualified.

Oct 6, 2026
Tone: positive
reported speed:
210.0 tokens/s generation
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentictool-uselong-contextcodingvision

User benchmarks Qwen3.8-27B NVFP4 at 210 tok/s single-stream on one RTX PRO 6000 Blackwell, using SGLang with the DFlash2 drafter. Setup is SGLang v0.5.20 with a 262,144-token window and 4 slots; the same model without a drafter ran 75 tok/s, and vLLM with DFlash2 ran 160 tok/s. The 27B matrix ran at a 400W power cap. User also reports 607 tok/s aggregate across 4 users, 97s full-window prefill, 73.3% BFCL core, 700/1000 arena score, and 24/40 CAPTCHA puzzles. Flash-Next and an uncensored 27B fine-tune were tested alongside.

Oct 3, 2026
reported speed:
63.7 tokens/s generation
quant:
NVFP4 (NVFP4)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports GLM-5.3-Flash at 63.7 tok/s single-request decode on 4x RTX PRO 6000 Blackwell GPUs. Setup is an experimental pinned SGLang build with an NVFP4 checkpoint, FP8 E4M3 KV cache, FlashInfer sparse MLA, and 262,144-token context, using tensor parallel TP4/EP4. With EAGLE speculative decoding (adaptive, up to 5 steps) the single-request figure rises to 82.8 tok/s. Eight concurrent requests give 183.8 tok/s aggregate baseline and 182.1 tok/s with EAGLE.

Sep 30, 2026
Tone: positive
reported speed:
77.6 tokens/s generation · 14-37 tokens/s prompt processing
quant:
BF16 (safetensors)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentictool-uselong-context

User reports Qwen3.8-27B at 77-80 tok/s on a single RTX PRO 6000 Blackwell 96GB. Setup is SGLang with BF16 safetensors, FP8 KV cache, EAGLE speculative decoding, and 262144 context length. The decode figure is a range; EAGLE accept rate was 0.94-1.00 and the user claims about 2.2x speedup over non-speculative decoding. Prefill was 14-37 tok/s on 122-token batches.

Sep 26, 2026
Tone: positive
reported speed:
11352.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-uselong-contextagenticcreative-writing

User reports Qwen3.8-Flash-Next on one RTX PRO 6000 Blackwell 96GB, with the official SGLang NVFP4 image cutting time to first token at roughly 254K context from 34.9s to 22.4s and raising prefill from 7,284 to 11,352 tok/s. Setup is the lmsysorg/sglang:dev-qwen38-next-local image with the RadixArk NVFP4 checkpoint, a 262,144-token window and one slot; the previous build streamed the 47.68 GiB per-layer embedding table from NVMe while the official image pins it in host RAM. A separate cache test dropped first-token time from 21.7s cold to 0.44s cached. Long-context recall scored 81/81, BFCL single-turn tool accuracy 85.2% versus 49.5% multi-turn, and tau2-bench telecom completion 68.1% with 41.2% passing all three attempts.

Sep 18, 2026

Qwen3.8 27B

2× NVIDIA RTX Pro 6000 Blackwell · SGLang · 262,144 ctx

Tone: mixed
reported speed:
150.0 tokens/s generation
quant:
FP8
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen 3.8 27B FP8 on two RTX 6000 Pro GPUs with sglang at 150 tk/sec, but tasks take 12x longer than Claude Opus 5. BF16 on the primary card with FP8 offloading took 3 hours versus 1h45m for FP8. Context lengths tested are 256k, 128k (too small) and 500k (faster despite being unsupported). A skill that reads large documentation files hits 80k context before starting. The user compares against Claude Code, Pi Code, Qwen Code and Hermes, and uses ChatGPT 5.6 high for testing.

Sep 7, 2026
Showing 1–6 of 6
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23