llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA Tesla P100 16GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
32.6 tokens/s generation · 493.0 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B Q6_K at 32.6 t/s decode (tg256, no MTP) on 2x Tesla P100 16GB with tensor split, versus 17.51 t/s upstream at the fork point. Setup is a llama.cpp fork with CUDA work for Pascal (sm_60), q4_0 KV cache, and fp16 math with fp32 accumulation. Prefill pp2048 at 0 context is 493 t/s versus ~250 t/s upstream. With MTP speculative decoding the fork reaches 54 t/s at 2k context and 29-35 t/s at 260k context. Prefill at 260k context is 123 t/s filling and 153 t/s for a question on a loaded context. Perplexity on the gate corpus at -c 4096 is 2.6101.

Oct 6, 2026
Tone: positive
reported speed:
110.0 tokens/s generation
quant:
Q4_0 (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User reports Nemotron-3.5-Lightning-30B-A3B at 110 tok/s writing code on two Tesla P100 16GB cards. Setup is llama.cpp b10970 with Q4_0 weights, F16 KV cache, tensor split across both cards, and the model's built-in MTP draft head at n-max 2. The same model reaches 95 tok/s on prose, 50 tok/s at 128k context, 39 tok/s at 256k, and 16 tok/s at 1M tokens. The study covers 545 speed measurements of 29 models from 2B to 122B parameters, all weights and KV cache in VRAM with no system RAM offload. A 119B MoE model runs 39 tok/s against 4.3 tok/s for a 70B dense model on the same cards. Tensor split makes dense models from 8B up 21-44% faster. A single P100 throttles to 906 MHz and loses 25% under sustained load, while two cards share the heat and lose 5.5%.

Oct 6, 2026
reported speed:
24.5 tokens/s generation · 448.7 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B Q4_K_XL at 24.5 tok/s generation and 448.7 tok/s prompt processing on two Tesla P100 16GB cards with tensor split. Setup is a patched llama.cpp (upstream b10660 plus eleven patches) with F16 KV cache and full GPU offload across two cards. The figures are the after-patch numbers from a patch series that improves decode and prefill; the same run measured 22.3 tok/s generation and 427.0 tok/s prompt processing before the patches. A real 7,655-token request reached 398.5 tok/s prompt processing, and four concurrent agents reached 20.3 tok/s each (72.9 aggregate).

Oct 3, 2026

Qwen3.8 27B

2× NVIDIA Tesla P100 16GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
50-60 tokens/s generation · 350.0 tokens/s prompt processing
quant:
Q6_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writing

User reports Qwen3.8 27B at 50-60 t/s generation and 350 t/s prefill on 2x Tesla P100 16GB with a custom llama.cpp fork. Setup is llama.cpp with Q6_K quant, 262144 context, batch 32768, ubatch 1024, and MTP speculative decoding. At 260k context, generation is 30-35 t/s and prefill is 110 t/s. GPUs are capped at 175W/250W each and run at 79C with minor thermal throttling; user estimates 5-10% higher numbers with better cooling.

Sep 27, 2026
Tone: positive
reported speed:
54-60 tokens/s generation · 440-500 tokens/s prompt processing
quant:
UD_Q4_K_XL (GGUF)
kv:
16bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writing

User reports Qwen3.6 35B A3B at 54-60 t/s generation (66-72 t/s on code) on a single Tesla P100 16GB, up from 30-35 t/s prose on an RX 6600 XT. Setup is llama.cpp with shinbunbun patches, UD_Q4_K_XL quant, 16-bit KV cache, MTP speculative decoding, and --n-cpu-moe 22, running 32k context. Prefill is 600 t/s at 0 ctx dropping to 440-500 t/s by 10k. The user notes the P100 is underrated for the price and has a second card coming for full offload.

Sep 27, 2026
Tone: mixed
reported speed:
85.0 tokens/s generation · 300.0 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports Qwen3.6-35B-A3B at ~85 t/s generation and ~300 t/s prompt processing on 2x Tesla P100. Setup is llama.cpp with the Q4_K_XL GGUF, MTP speculative decoding, and community P100 patches that added about 50% decode; a single P100 on Q2_K_XL reached ~76 t/s generation and ~250 t/s prompt processing. Card count barely affects single-stream speed, and PCIe lane width (x16/x16, x16/x8, x8/x8) made no difference. The user corrects an earlier 3-card result that was invalidated by stuck 405 MHz core clocks.

Sep 24, 2026
Tone: positive
reported speed:
70.0 tokens/s generation
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports almost 70 t/s with Qwen3.6-35B-A3B Q4_K_XL on a budget build using three Nvidia Tesla P100 16GB GPUs. The GPUs cost $80 each and are split across three nodes to manage thermals; the motherboard required a patched BIOS to enable Above 4G Decoding. User is still testing and optimizing, and has published a GitHub repository for the build.

Sep 22, 2026
Tone: positive
reported speed:
40.0 tokens/s generation
quant:
Q6_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 40+ t/s on two Tesla P100s, with a roughly 30 t/s average across long contexts and up to 55 t/s at 0 context. Setup is a custom llama.cpp fork with P100 kernel optimizations, Q6_K quant, both cards capped at 175W and communicating over PCIe gen 3. Speeds vary by about ±2 t/s; the user notes fp16 math saves 40% or more on prefill with negligible accuracy loss.

Sep 20, 2026
Showing 1–8 of 8
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423