llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX 5060 Ti 16GB
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Qwen3.8 27B

2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 52,000 ctx

Tone: mixed
reported speed:
22-23 tokens/s generation
quant:
UD-Q6_K_XL (GGUF)
kv:
f16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User reports Qwen3.8-27B GGUF quants fail to converge on long thinking tasks on 2x RTX 5060 Ti 16GB, while NVFP4 mixed-precision quants converge in ~17K thinking tokens. Setup is llama.cpp and vLLM with UD-Q6_K_XL GGUF (23.6 GB) and f16 KV cache, tensor-parallel across 2 GPUs. Throughput was healthy at 23-41 t/s on llama.cpp and 22-23 t/s on vLLM. The failure appears after 30K-38K thinking tokens with no closing token, tail-looping, or engine crash. The same model in NVFP4 (FP8 attention/GDN, FP4 MLP) converges in ~17K tokens on SGLang at 47-50 t/s. User hypothesizes recurrent state error accumulation under uniform quantization.

Oct 6, 2026
Tone: positive
reported speed:
95.2 tokens/s generation · 3476.0 tokens/s prompt processing
quant:
MXFP4-MOE (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma 4 26B-A4B-it at 95.22 tok/s decode and 3,476 tok/s prompt processing on an RTX 5060 Ti 16 GB. Setup is a custom llama.cpp build (commit 0b484ab2b plus 12 commits) with MXFP4-MOE quantization (experts MXFP4, dense Q8_0, 4.66 BPW) and q4_0 KV cache, 1 slot, flash attention enabled, 65,536 token context. The MXFP4-MOE quant brings the model to 13.70 GiB, fitting the 16 GB card where Q4_K_M at 15.85 GiB does not. On an RTX 5090 the same quant gives 10,733 tok/s prompt and 196.6 tok/s decode versus 8,744 and 219.9 for Q4_K_M, with perplexity 3,864 vs 3,615.

Oct 6, 2026
reported speed:
30.5 tokens/s generation · 250.4 tokens/s prompt processing
quant:
IQ1_M (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8-Flash-Next at 30.5 t/s generation and 250.4 t/s prompt processing on 2x RTX 5060 Ti with 32GB system RAM. Setup is llama.cpp with an IQ1_M GGUF (27.58 GB), 98304 context, q8_0 KV cache, tensor split 1,1, and 8 MoE layers offloaded to CPU. User asks whether their settings are correct and what config others run for this model on this hardware.

Oct 5, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 4,096 ctx

reported speed:
29.2 tokens/s generation
quant:
UD-IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.8 27B at 29.18 tok/s decode on an RTX 5060 Ti 16 GB at 4K context. Setup is llama.cpp build ad1de39e0 with UD-IQ3_XXS quant and q4_0 K/V cache, full GPU offload, flash attention, one parallel slot. Baseline decode without speculation is 29.18 tok/s at 4K, 23.62 at 32K and 19.77 at 64K; MTP n=2 gives 56.91/42.53/34.95 and MTP n=3 gives 63.89/45.48/39.06 tok/s at the same contexts. N-gram speculation added only 1.7% (29.66 tok/s). A 96K prompt measured 617.73 prompt tok/s and 36.74 decode tok/s.

Oct 3, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 94,208 ctx

reported speed:
50-55 tokens/s generation
quant:
UD-IQ3_XXS (GGUF)
kv:
Q4_0
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at a steady 50-55 t/s with MTP on a single RTX 5060 Ti 16GB. Setup is llama.cpp with UD-IQ3_XXS GGUF weights and Q4_0 KV cache for K and V, 94208-token context, batch and ubatch 512, flash attention on, draft-mtp speculation with spec-draft-n-max 3, parallel 1. Without MTP the average is 35 t/s. The machine has an Intel i5 10th gen and 32GB DDR4.

Oct 3, 2026
Tone: mixed
reported speed:
35.0 tokens/s generation
quant:
IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports about 35 t/s or more with Qwen3.8 27B GSQ-RCO-Uncensored on an RTX 5060 Ti 16GB. Setup is llama.cpp with IQ3_XXS quant, q4_0 KV cache, 131072 context, and MTP draft plus ngram speculative decoding. The user is on Fedora 44 with an AMD Ryzen 9600x and 16 GB system RAM, and had struggled to get reasonable speed before finding this model.

Oct 1, 2026
reported speed:
47-51 tokens/s generation · 1168.0 tokens/s prompt processing
quant:
UD-IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.5-35B-A3B at 47-51 tok/s generation on an RTX 5060 Ti 16GB over OCuLink in a Proxmox VM. Setup is llama.cpp with UD-IQ3_XXS quant and q4_0 KV cache, fully on the GPU with 348 MiB headroom, at 160K context. Prompt eval of 75K tokens took 64.8 s (1,168 tok/s). The 9B variant reached 40-50 tok/s with UD-Q4_K_XL and about 47-51 tok/s with IQ3_XXS.

Sep 30, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
30-40 tokens/s generation
quant:
IQ3_S (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentictool-uselong-contextvisioncoding

User reports Qwen3.8-27B at 256K context on an RTX 5060 Ti 16GB, with decode staying around 30-40 tok/s. Setup is a KVMem-enabled llama.cpp server (v0.14.0) with ISTA IQ3_S quant, MTP, Q8 KV cache, vision on GPU, and a 32K GPU window plus 16K reserved for new tokens; the rest of the KV cache stays in 32GB system RAM. In a 33-request tool task ending at 262058/262144 tokens, prefill was 437 tok/s first pass and 243 tok/s overall, decode 30 tok/s over the tool rounds and 29 tok/s on the last 512 tokens, with 15.5 GB peak VRAM and 13.1 GB RAM. An IQ4_XS variant with Q5 KV and CPU vision reached 466/255 tok/s prefill and 33/35 tok/s decode. A shorter image and code task ran about 37-41 tok/s decode.

Sep 29, 2026
Tone: positive
reported speed:
80.0 tokens/s generation · 4737.5 tokens/s prompt processing
quant:
NVFP4 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Nemotron 3.5 Lightning 30B-A3B at 79.98 t/s generation and 4737.46 t/s prompt processing on 2x RTX 5060 Ti 16GB. Setup is llama.cpp with NVFP4 GGUF weights and q8_0 KV cache, 1048576 token context, flash attention on, no MTP layers. The run processed a 54025-token prompt and generated 2156 tokens; the user notes it fits in 32 GB VRAM with no expert-layer offloading to RAM.

Sep 28, 2026

Unknown family

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 98,304 ctx

Tone: mixed
reported speed:
20-35 tokens/s generation
kv:
kvarn3
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports failing to reproduce a claimed 50+ t/s on an RTX 5060 Ti 16GB, getting 20-35 t/s instead. Setup is a freshly built llama.cpp fork (beellama.cpp) for sm120 with MTP speculative decoding, kvarn3 KV cache, 98304 context, and parallel 1. The user questions whether the original high-throughput post is fake or if they are missing something.

Sep 28, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 100,000 ctx

Tone: positive
reported speed:
25-30 tokens/s generation
quant:
IQ3_XXS (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionlong-context

User reports Qwen3.8 27B running on an RTX 5060 Ti 16GB with MTP speculative decoding, reaching up to 50 t/s in the first half of a 100k context and dropping to about 25-30 t/s at the end. Setup is a modified llama.cpp fork with adaptive KV cache streaming, IQ3_XXS GGUF weights, q8_0 K cache and q4_0 V cache, a 2100 MiB KV stream buffer, and a BF16 vision head. The user notes the streaming fork only works on NVIDIA hardware and that the speed depends on context size versus KV stream buffer size.

Sep 28, 2026
reported speed:
41.6 tokens/s generation · 1718.3 tokens/s prompt processing
quant:
Q8_0 (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports MiMo 2.6 Distill Qwen 9B at 41.63 t/s generation and 1718.29 t/s prompt processing on an RTX 5060 Ti 16GB. Setup is llama.cpp (Llama UI) with GGUF Q8_0 weights, q4_0 KV cache, 122880-token context, and Flash Attention enabled. The run produced 11979 output tokens over 4 min 47 s from a 1929-token prompt; the first generated version failed and a second review pass by the model produced the final working demo.

Sep 27, 2026

Qwen3.8 27B Swift-Genesis

2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx

Tone: mixed
reported speed:
76.0 tokens/s generation
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionagenticlong-context

User reports Qwen3.8 27B Swift-Genesis at 76 t/s in their benchmark on 2x RTX 5060 Ti 16GB, with a range of 40-113 t/s and 45-65 t/s under normal agentic work. Setup is llama.cpp with GGUF weights at 262k context, MTP4 speculative decoding, and vision offloaded to system RAM. The user notes the model fits under 17GB with vision, allowing full 262k context on 32GB VRAM, and asks what the tradeoff of this model is compared to other Swift NVFP4 models.

Sep 27, 2026

Occamy 1.0

2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 175,000 ctx

Tone: positive
reported speed:
100.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Occamy 1.0 at around 100 tok/s on two RTX 5060 Ti cards, with enough memory for two independent ~175K context pools for concurrent subagents. Setup uses llama.cpp on the desktop worker box; the cards cost about $400 each. A 96 GB M5 Ultra Mac Studio is planned as the primary inference appliance running Qwen 3.8 Next Flash through oMLX, with DeepSeek V4.1 Flash via API filling in until it arrives.

Sep 24, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 110,000 ctx

reported speed:
5.3 tokens/s generation
quant:
IQ4 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Qwen3.8 27B at ~5.3 t/s on an RTX 5060 Ti 16GB with 16GB single-channel system RAM. Setup is llama.cpp with IQ4 weights, 110k context, Q8 KV cache, and partial GPU/CPU offload; CPU and GPU each sit around 50% utilization. User estimates Q8 weights plus 256k context would need 38-40GB total and drop to roughly 2 t/s, and asks whether quant level, context, or throughput should be prioritized for a local coding-agent worker.

Sep 17, 2026
Tone: mixed
reported speed:
18.3 tokens/s generation · 25.4 tokens/s prompt processing
quant:
Q4 (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next-REAP-320 at 18.3 t/s generation and 25.4 t/s prefill on an RTX 5060 Ti 16GB with 32GB system RAM. Setup is llama.cpp with a Q4 GGUF, 64k context, q4_0 KV cache, --n-cpu-moe 34, --ngl 48, and lazy mmap for the 29.48 GB n-gram embedding. User notes the model is smarter than Qwen3.8 27B but keeps the 27B as daily driver due to 5x faster prompt processing.

Sep 14, 2026
Tone: positive
reported speed:
37.2 tokens/s generation
mtp (multi-token prediction):
off

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User benchmarks multiple models on 16 GB of VRAM. Qwen3.6-35B-A3B-APEX-I-Quality reaches 96.7% HumanEval pass@1, 100% agentic success and 37.2 tok/s. Ornith-1.0-35B IQ4_NL runs at 38 tok/s. Other models tested are Qwen3.8-27B variants, gemma-4-26B-A4B, KAT-Coder-V2.5-Dev and Qwen3.5-9B. Ornith 1.0 35B A3B wins overall.

Sep 9, 2026

Muse 30B Glimmer

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 131,768 ctx

reported speed:
18.0 tokens/s generation
quant:
Q4_K_XL (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a model fits on a single RTX 5060 Ti 16GB at 131k context with a Q4 KV cache. Setup uses GGUF weights only, with no dflash or mmproj loaded. A Q8 KV cache allows about 90k context.

Sep 7, 2026
Tone: positive
reported speed:
40.0 tokens/s generation
quant:
Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

vision

User reports Qwen3.6-35B-A3B MoE at about 40 t/s on an RTX 5060 Ti 16GB with mmproj enabled, and just over 50 t/s without mmproj and with --n-cpu-moe 24. Setup is llama.cpp server in Docker with --cache-type-k/v q8_0, --cache-ram 0, --no-mmap, and --n-cpu-moe 19, or 24 without vision. The user experienced an OOM shutdown before adding --cache-ram 0 and is seeking advice on the setup.

Sep 7, 2026
reported speed:
11.0 tokens/s generation · 200.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports llama.cpp running on 4x RTX 5060 Ti 16GB with DDR4 3200 RAM in 4-channel. Setup uses -ub/-b at 4096.

Sep 7, 2026
Showing 1–20 of 29
Page 1 of 2

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23