llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX 5090
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
287.0 tokens/s generation · 13700.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User compares NInfer's official Qwen3.8-27B quant (part NVFP4, part FP8) against QUASAR's QAT full-NVFP4 checkpoint with DFlash2 embedded, both on a single RTX 5090. QUASAR reaches 484 tok/s on JSON output, 204 tok/s on prose, and 13.7k tok/s prefill, using 27.0 GB VRAM. The official quant gets 400 tok/s JSON, 176 tok/s prose, 11.4k tok/s prefill, and 30.8 GB VRAM. On a real agentic task the QUASAR build averages 287 tok/s over 33k tokens. GPQA Diamond scores 89.9% versus 90.4% for the official quant, and perplexity is about 1.9% worse.

Oct 3, 2026

Qwen3.8 27B

NVIDIA RTX 5090 · NInfer · 240,000 ctx

Tone: positive
reported speed:
158.0 tokens/s generation · 7265.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextsummarizationtool-useagentic

User reports Qwen3.8-27B at 158 tok/s decode and 7,265 tok/s prefill on an RTX 5090 32GB (eGPU via OCuLink Gen4 x4). Setup is NInfer with NVFP4 weights, FP8 KV cache, 240K context, MTP3 speculative decoding at 76% acceptance, and 2 concurrent lanes. Decode rises to 213 tok/s at 32K and 202 tok/s at 128K; prefill is 6,892 tok/s at 32K and 3,904 tok/s at 128K. User compares against llama.cpp (Q5_K_M GGUF, q8_0 KV, 196K context, MTP on) at 114 tok/s decode and 1,545 tok/s prefill at 1K, and notes vLLM at about 70 tok/s decode at short context. Quality was statistically indistinguishable across engines on a 250+ item eval. NInfer does not support json_mode.

Sep 27, 2026

Qwen3.8 27B Swift-1.5

NVIDIA RTX 5090 · NInfer · 262,144 ctx

reported speed:
160.8 tokens/s generation · 3269.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
k8v4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-contextvision

User reports Swift-1.5 Qwen3.8-27B at 160.8 t/s decode on a single RTX 5090 at 262,144 context. Setup is the NInfer v3 engine with an all-NVFP4 (W4A4 gs16) artifact plus a z-lab DFlash2 drafter at K=7, k8v4 KV cache, 18.0 GiB of weights in VRAM, 450 W power cap, and concurrency 4. Prefill at 200k context measured 3,269 t/s. IFBench prompt-strict 69.0, prompt-loose 72.7, instr-strict 70.4, instr-loose 73.6; GSM8K-200 95.0% (190/200); long-context needle at 250,031 tokens exact across 3 depths.

Sep 25, 2026

Qwen3.8 27B Uncensored

NVIDIA RTX 5090 · NInfer · 262,144 ctx

Tone: positive
reported speed:
175.0 tokens/s generation
quant:
NVFP4
kv:
Q4
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcodinglong-context

User reports Qwen3.8 27B uncensored at 175 t/s on an RTX 5090 32GB. Setup is NInfer with NVFP4 / groupwise-int quantization and MTP enabled, running at 262,144 context with a Q4 KV cache. The user says the Q4 KV cache matched Q8 quality and that the longer context enabled long reasoning tasks in a Hermes harness.

Sep 24, 2026
Tone: positive
agentictool-use

User reports a native ninfer implementation of a Jev-like decision API running Qwen3.8 27B (Swift fine-tune) on an RTX 5090, scoring 84.4% (195/231) on the JevBench v1.2 public set with three in-flight decisions and a 19 s wall clock. Setup drives the native POST /v1/decisions route directly rather than the chat-completions path; latency p50 is 127 ms and p95 is 699 ms, with hard-tier p95 prefill-dominated by ~3.7k-token states and tail latency inflated by a single serialized decision worker. Calibration is ECE 0.045 and Brier 0.081; the user calls it a dirty POC with bugs and incorrect allocation still to fix, and notes latency is not great but accuracy is almost on par.

Sep 23, 2026
Tone: positive
reported speed:
178.0 tokens/s generation
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B NVFP4 at 178 t/s on an RTX 5090 after undervolting the GPU and capping CPU power at 50%. Setup uses the Ninfer inference engine. GPU memory was overclocked by 2400 MHz and the GPU was undervolted; CPU power was limited to 50% with no measurable TPS loss. A systemd service polls GPU usage and applies the CPU cap automatically during inference. Stock settings gave 172 t/s in synthetic tests; undervolting and VRAM overclocking raised this to 178 t/s, and DFlash2 speculative decoding currently yields 200 t/s. GPU power dropped from 600 W to under 450 W and GPU temperature from 75°C to 62°C.

Sep 18, 2026
Tone: positive
reported speed:
205.0 tokens/s generation · 9000.0 tokens/s prompt processing
quant:
NVFP4 (NVFP4)
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writinglong-contextvision

User reports Qwen3.8 27B Huihui abliterated NVFP4 at 205 t/s decode on a single RTX 5090 32GB, with prefill around 9,000 t/s at 8K context. Setup is NInfer on Windows 11 + WSL2 with fp8 KV cache at 196,608 context, MTP speculative decoding with 3 draft tokens and --lm-head-draft, vision enabled, weights about 19.7 GB. Decode varies with MTP acceptance: 253 t/s on predictable text, 205 t/s on coding, about 120 t/s on creative prose, and 60-80 t/s with MTP off. At 128K context decode is about 115 t/s and prefill about 4,500 t/s.

Sep 17, 2026

Qwen3.8 27B

NVIDIA RTX 5090 · NInfer · 200,000 ctx

Tone: positive
reported speed:
162.1 tokens/s generation · 4260.0 tokens/s prompt processing
quant:
Q8
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at ~162.1 t/s generation and ~4.26k t/s prefill on an RTX 5090. Setup is the Ninfer Windows Edition with a Q8 KV cache at 200k context. Switching from LM Studio Q6 to Ninfer roughly doubled both decode and prompt speeds.

Sep 17, 2026

Qwen3.8 27B

NVIDIA RTX 5090 · NInfer · 555,000 ctx

Tone: positive
reported speed:
117.0 tokens/s generation · 1600.0 tokens/s prompt processing
kv:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagenticcoding

User reports a fork of NInfer with a custom NVFP4 KV cache, YaRN context extension to 555k, multi-level prefix reuse with a host KV safety net, tool calling improvements, and monitoring. Benchmarks show decode at 117 tok/s at 400k+ context, cold prefill of 260s for 414k tokens at 1600 tok/s, H2D restore in 0.4s, and 30GB of host KV. Quality results show LongBench matching int8, AIME at 96.7%, and needle-in-haystack at 100%.

Sep 9, 2026
Tone: positive
reported speed:
200.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at about 200 t/s generation on a single RTX 5090 with speculative decoding, on day-0 support in NInfer. The setup uses a shared paged KV cache and ReplaySSM for GDN, with PDL in use. The user notes support for up to 8 concurrent requests.

Sep 7, 2026

Qwen3.8 27B

NVIDIA RTX 5090 · NInfer · 262,144 ctx

Tone: positive
reported speed:
200.0 tokens/s generation · 5950.0 tokens/s prompt processing
quant:
NVFP4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports an aggregate 880 t/s at 6 parallel requests, peaking at 967 t/s, and over 200 t/s single stream with MTP speculative decoding. Prefill runs at about 5,950 t/s. Weights take 16.8 GiB, leaving about 13 GiB for KV cache. Benchmarks are HumanEval+ 152/164 and AIME25+AIME26 55/60, identical to int4 but 1.56x-1.98x faster. The run requires Blackwell FP4 cores and a 6-line patch not yet upstream.

Sep 7, 2026
reported speed:
169.7 tokens/s generation
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports median decode speeds by context size on an unspecified setup: 169.7 t/s under 50k, 167.6 t/s at 50-100k, 156.9 t/s at 100-150k, 149.0 t/s at 150-200k, and 143.7 t/s at 200k and above. Context length is set to maximum. Peak decode for a single request is 222.0 t/s, with 211.2 t/s sustained over 5s.

Sep 7, 2026

Qwen3.8 27B

NVIDIA RTX 5090 · NInfer · 240,000 ctx

Tone: positive
reported speed:
170-179 tokens/s generation
quant:
NVFP4
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B averaging in the 170s t/s on an RTX 5090, peaking at 220 t/s. Setup is ninfer-serve with the nvfp4 weights, MTP speculative decoding with 3 draft tokens, lm-head-draft, an fp8 KV cache and 240,000 context. User calls it double or more the throughput of llama.cpp.

Sep 7, 2026

Qwen3.8 27B

NVIDIA RTX 5090 · NInfer · 240,000 ctx

Tone: positive
reported speed:
202.0 tokens/s generation · 3904.0 tokens/s prompt processing
quant:
NVFP4
kv:
FP8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contexttool-useagentic

User compares NInfer, llama.cpp and vLLM for Qwen3.8-27B on an RTX 5090. NInfer with NVFP4 and MTP3 reaches 158-213 t/s decode and 3904-7265 t/s prefill. NInfer outperforms llama.cpp with Q5_K_M by 1.4-2.8x decode and 2.6-4.7x prefill. Quality is statistically indistinguishable across engines. NInfer lacks json_mode support, and vLLM speed is not directly comparable due to wall-clock timing.

Sep 7, 2026
Tone: positive
reported speed:
200.5 tokens/s generation · 2587.0 tokens/s prompt processing
quant:
NVFP4
kv:
Q8
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User benchmarks 180K context with MTP5 enabled, reaching 2587 t/s prefill and 200.5 t/s decode. The run also covers 64K and 120K contexts. A synthetic corpus yields high MTP acceptance, while real agentic coding logs show about 51% accept and about 154 t/s.

Sep 7, 2026

Qwen3.8 27B

2× NVIDIA RTX 5090 · NInfer · 1,048,576 ctx

Tone: positive
reported speed:
119.0 tokens/s generation
quant:
NVFP4
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a fork of NInfer with tensor parallelism and YaRN scaling decoding at 119 t/s with MTP and 57 t/s without at 653k context, and 48 t/s decode at 1M context, around 100 t/s with MTP. Setup is a fork of NInfer with tensor parallelism and YaRN scaling. Against vLLM at 42 t/s at 653k, MTP acceptance drops to zero past 262k. Prefill is 1.2-1.3x slower than vLLM. Two GPUs reach 75 t/s versus 54 t/s on one at 250k.

Sep 7, 2026
Tone: positive
quant:
NVFP4

User compares NVFP4 against Q5_K_M on Qwen3.8-27B. NVFP4 with adjusted sampling (temp=0.9, min_p=0.05) reaches 80% strict on IFBench, matching local BF16 and near official BF16 79.5%, while Q5_K_M stays at 76%. NVFP4 runs about 3x faster and uses less VRAM. The Q5_K_M run uses llama.cpp.

Sep 7, 2026
Showing 1–17 of 17
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23