llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX 2080 Ti 22GB (modded)
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
69.0 tokens/s generation
quant:
UD-IQ4_XS (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.6-35B-A3B at 4-bit (UD-IQ4_XS) scoring 89.63% pass@1 on HumanEval at 69 tok/s on a single RTX 2080 Ti 22GB. Setup is llama.cpp with UD-IQ4_XS dynamic 4-bit (4.25 bpw), q8_0 KV cache, 16384 context, flash-attn on, whole model resident in VRAM with no offload. A routing patch (MoE expansion, 20 experts instead of 8 on layers 25-39) reached 90.85% pass@1 at 56 tok/s, a 19% decode speed cost; the user calls the +2 problems within statistical noise and notes the tests are original HumanEval, not EvalPlus.

Oct 7, 2026
Tone: positive
reported speed:
59-68 tokens/s generation
quant:
GPTQ-Int4 (GPTQ)
kv:
turboquant_k3v4_nc

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentictool-usevisionlong-context

User reports Qwen3.8-27B at ~59-68 tok/s single-stream decode on 2x modded RTX 2080 Ti 22GB (SM75, TP=2, NVLink). Setup is a vLLM fork with GPTQ-Int4 self-quant, turboquant_k3v4_nc KV cache, 524,288 max-model-len (YaRN 2x), MTP K=2 speculative decoding, and vision enabled. Aggregate throughput is 364.7 tok/s at 16 lanes; 24 lanes regresses. The S4 scoped re-emission drafter measured 811.5 tok/s on copy-shaped spans but crashes at current HEAD, and v3 GPU-merge async is negative.

Oct 7, 2026
Tone: positive
reported speed:
95.5 tokens/s generation · 465.8 tokens/s prompt processing
quant:
Q4_K_P (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen3.6-35B-A3B at 95.5 t/s decode and 465.8 t/s prefill on a mixed three-GPU Turing setup of 2x CMP 50HX 10GB plus 1x RTX 2080 Ti 22GB. Setup is llama.cpp with a ported DP2A patch (PR #25834) and -fmad=false, Q4_K_P GGUF weights, Q8_0 KV cache, Flash Attention, 262144 context, MTP speculative decoding at --spec-draft-n-max 3, tensor split 1,1,2.5 with the RTX as tail stage, ubatch 448. Baseline DP4A gave 47.8 t/s decode and 372.7 t/s prefill; DP2A alone 54.0 t/s; adding -fmad=false 62.2 t/s. A repo-28k workload measured 405.4 t/s prefill and 82.3 t/s decode. An end-to-end rerun showed 91.4 t/s raw eval, 4% below the 95.5 reference. Context 368640 also works but quality beyond the trained window was not measured.

Oct 7, 2026
Tone: positive
reported speed:
36.6 tokens/s generation
quant:
IQ3_S (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagentictool-usevision

User reports Qwen3.8-27B at 36.6 tok/s decode on a single modded RTX 2080 Ti 22GB, with 262,144 tokens of context. Setup is KVMem (retrieval-based long context) with IQ3_S weights, q8_0 KV cache, MTP speculative decoding (58.6%/61.6% acceptance) and vision, using 16,552 of 22,528 MiB VRAM. The user compares against the upstream author's RTX 5060 Ti 16GB at 31.7 tok/s, and notes the lossless KV-streaming alternative runs 30-42 tok/s in the resident window but drops to ~9.6 tok/s past 135K tokens; a 260,096-token needle-in-a-haystack test hit exactly.

Oct 7, 2026
Tone: positive
reported speed:
101.3 tokens/s generation · 1841.7 tokens/s prompt processing
quant:
AWQ (AWQ)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen3.6-27B-AWQ at 101.3 tok/s decode and 1841.7 tok/s prefill on dual modified RTX 2080 Ti 22GB cards with NVLink. Setup is vLLM 0.21.0 with AWQ Marlin, TP=2, MTP K=3, FlashInfer/FA2 attention, and FlashQLA SM70/SM75 legacy GDN prefill. The speed columns use the PP4096/TG128 repeat. The same rig reached a 735,084 token KV cache with turboquant_4bit_nc at max_model_len=262144, and passed a PP262000/TG1 gate at 785.26 tok/s prefill. A sequential 60-request Ragent6 run averaged 700.9 tok/s prefill and 35.2 tok/s generation. Gemma4 31B GPTQ reached 99.64 tok/s decode and 1655.65 tok/s prefill on the same runtime.

Oct 7, 2026
Tone: positive
reported speed:
44.1 tokens/s generation
quant:
IQ3_S (GGUF)
kv:
Q8_0
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextvision

User reports Qwen3.8-27B-GSQ-RCO at 44.07 t/s decode on a hardware-modded RTX 2080 Ti 22GB, with 38~44 t/s quoted as the overall range. Setup is a llama.cpp fork (TurboQuant 4-bit + Adaptive KV Streaming) with IQ3_S weights, 256K context (262,144 tokens), Q8_0 key cache and turbo4 value cache, a 2048 MiB GPU staging pool, and MTP speculative decoding with 2 draft tokens. Needle tests at 8K/16K/32K gave 41.23, 39.00 and 34.84 t/s decode with 81-83% draft acceptance and 100% recall at 82% depth; a sustained 256K run gave 40.55 t/s at 89.17% acceptance. Peak VRAM was 18,619 MiB and GPU power peaked at 266.4 W.

Sep 28, 2026
Tone: positive
reported speed:
25.0 tokens/s generation
quant:
W8A16
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at ~25 tok/s with standard autoregressive decoding (MTP0) on a modded RTX 2080 Ti 22GB. Setup is a Turing port of the NInfer engine with the official groupwise-int (W8A16) artifact and Q8 KV cache. With MTP3 speculative decoding (draft window = 3) the title gives 45 tok/s and the body ~456 tok/s, at ~65% acceptance rate. VRAM usage is ~17.5 GiB with the MTP draft weights loaded, leaving ~4.5-5.0 GiB free for KV cache.

Sep 7, 2026
Showing 1–7 of 7
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423