llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA CMP 170HX 64GB (unlocked)
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
117.0 tokens/s generation · 6066.0 tokens/s prompt processing
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4.1-Flash at 117 tok/s decode on one stream at 128k context and 6,066 tok/s prefill at 105k tokens on eight 64 GB CMP 170HX mining cards. Setup is vLLM with fp8 KV cache, PP=8, DSpark speculative decoding with 5 draft tokens, Engram tables in pinned host RAM, and 1M context. Decode drops to 96 tok/s at 512k. Eight concurrent streams reach 532 tok/s aggregate (66 per stream) at 105k. KV pool holds 6.17M tokens. One card was capped to 180 W after PCIe bus drops.

Oct 5, 2026
Tone: positive
reported speed:
101.3 tokens/s generation · 3191.0 tokens/s prompt processing
quant:
W4A16 (W4A16)
kv:
BF16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-contexttool-usemultilingual

User reports Qwen3.8-Flash-Next at 101.3 t/s decode on Japanese prose and 170.0 t/s on code, single stream, on two NVIDIA CMP 170HX cards. Setup is vLLM with W4A16 weights, BF16 KV cache, MTP k=4 speculative decoding, 262,144-token context, and expert parallel across the two cards. The PLE n-gram table was converted to FP8 locally to fit in 92 GiB of host RAM. Aggregate throughput reaches 392 t/s at 4 concurrent requests. Prefill measures 3,191 t/s at 6,954 tokens and 3,261 t/s at 27,853 tokens. A 200,087-token prompt returned the planted code with 70.5 s time to first token and 69.0 t/s decode at depth. MTP is worth about 1.6x on prose and 2.6x on code. The user notes xhigh reasoning effort spent the whole budget without answering in 5 of 6 runs.

Sep 27, 2026
reported speed:
95.0 tokens/s generation
quant:
FP8
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User reports DeepSeek V4 Flash Vision at up to ~95 tok/s on 4× CMP 170HX 64GB (256GB aggregate HBM2e). Setup is vLLM with pipeline parallel 4, max_model_len 262144, FP8 KV cache, DSpark speculative decoding with 6 tokens, and max_num_seqs 2. Normal coding workload runs ~50–70 tok/s, strong speculative-decoding periods ~80–90 tok/s, peak ~95 tok/s. Cards sit around 56–62GB VRAM each at 50–60°C. Next tests planned for Qwen3.8 Flash Next and GLM 5.3.

Sep 24, 2026
reported speed:
116.0 tokens/s generation · 6800-7950 tokens/s prompt processing
quant:
W4A16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3-Next-80B at 116 t/s single-stream decode on a single NVIDIA CMP 170HX with 64 GB HBM2e, at a 150 W cap. Setup is vLLM 0.27.1 with W4A16 weights (40.9 GB), torch 2.13.0+cu130, CUDA 13.0, Ubuntu 26.04 LTS, on an AMD Ryzen Threadripper PRO 3945WX with 128 GB DDR4 ECC. Prefill over about 8.9k tokens measured 6800 to 7950 t/s; power draw 137 to 145 W. Aggregate throughput at 8 concurrent requests was 352 t/s, saturating at 4 slots. Also measured on the same card for context: Ornith-1.5-35B FP8 at 122.5 t/s and Qwen3.8-27B W4A16 with DFlash2 at 127 t/s single-stream.

Sep 23, 2026
reported speed:
98.1 tokens/s generation · 5300.0 tokens/s prompt processing
quant:
MXFP4+FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731 at 98.1 t/s single-stream decode on 4x CMP 170HX (64 GB HBM2e each), up from a 50.8 t/s baseline with DSpark speculative decoding. Setup is vLLM with pipeline parallelism, native MXFP4+FP8 weights (~155.4 GiB), context verified to 1,047,736 tokens. At 100k context single-stream decode is 38.8 t/s with 14.6 s time to first token. Aggregate throughput across 64 concurrent requests is 712.8 t/s with DSpark (472.0 t/s baseline), and 90.0 t/s at 100k context. Prefill across the 25k-77k context range is about 5,300 t/s.

Sep 23, 2026
Tone: positive
reported speed:
147.0 tokens/s generation
quant:
W4A16 (W4A16)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-context

User reports Qwen3.8-27B at 147.0 t/s single-stream decode on one NVIDIA CMP 170HX 64GB, with 134.7 t/s at 4K, 100.1 t/s at 65K, ~90 t/s at 126K, and 64.9 t/s at 250K context. Setup is vLLM 0.27.1 with a W4A16 target, a DFlash2 W4A16 drafter, FP8 target KV, BF16 draft KV, a custom SM80 split-KV verifier, full CUDA Graph, 35 verifier segments / 140 CTAs, k=3 draft tokens, MAX_SEQS=1, and 1350 MHz / 180W. These are decode-only numbers, not end-to-end throughput including prefill. FP8 beat INT8 at long context (48.3 vs 115.4 ms/iter at 250K). Allowing 2 concurrent long requests made makespan and slowest-request throughput worse at both 126K and 250K.

Sep 20, 2026
Tone: mixed
reported speed:
50.0 tokens/s generation
quant:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at around 50 t/s on a CMP 170HX, running at Q8 or BF16. The user finds the model hallucinates classes and APIs in 90%+ of answers to domain-specific enterprise Java questions, while DeepSeek and Microsoft Copilot produce working code 99% of the time. The user asks whether Qwen's agentic optimization is responsible and requests tips for a harness, system prompt, or skills to reduce hallucination.

Sep 18, 2026
Tone: positive
reported speed:
63.0 tokens/s generation · 1700.0 tokens/s prompt processing
quant:
Q6_K
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B at 63 t/s on 4x CMP 170HX 8GB cards flashed to 64GB each, 256GB total. Setup uses MTP with little-MoE default. Generation reaches 110 t/s with MTP optimistic.

Sep 7, 2026
reported speed:
29.0 tokens/s generation · 450.0 tokens/s prompt processing
quant:
Q4_K_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports testing multiple models on 4x CMP 170HX 64GB cards, including DeepSeek V4-Flash 0731 with a Q4_K_XL quant, 13B active MoE, and 1M context, plain no-spec. The cards are cut-down A100 mining cards with 64GB each, PCIe Gen2 x4, and no NVLink. Other models tested include gpt-oss-120B, Qwen3.6-35B-A3B, GLM-4.5-Air, and MiniMax-M2.7.

Sep 7, 2026
Tone: positive
reported speed:
3468.0 tokens/s prompt processing
quant:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports an unlocked CMP 170HX, A100 silicon with a firmware unlock, reaching 193 TFLOPS tensor throughput, up from 6.3 TFLOPS. In llama.cpp, pp512 rises from 599.6 to 3468 tok/s. Serving Qwen3.8-27B-FP8 under vLLM, one 170HX beats a 2x3090 tensor-parallel pair on prefill at 197W versus 454W.

Aug 28, 2026
Showing 1–10 of 10
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423