llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: NVIDIA RTX 5060 Ti 16GB
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
95.2 tokens/s generation · 3476.0 tokens/s prompt processing
quant:
MXFP4-MOE (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma 4 26B-A4B-it at 95.22 tok/s decode and 3,476 tok/s prompt processing on an RTX 5060 Ti 16 GB. Setup is a custom llama.cpp build (commit 0b484ab2b plus 12 commits) with MXFP4-MOE quantization (experts MXFP4, dense Q8_0, 4.66 BPW) and q4_0 KV cache, 1 slot, flash attention enabled, 65,536 token context. The MXFP4-MOE quant brings the model to 13.70 GiB, fitting the 16 GB card where Q4_K_M at 15.85 GiB does not. On an RTX 5090 the same quant gives 10,733 tok/s prompt and 196.6 tok/s decode versus 8,744 and 219.9 for Q4_K_M, with perplexity 3,864 vs 3,615.

Oct 6, 2026
reported speed:
30.5 tokens/s generation · 250.4 tokens/s prompt processing
quant:
IQ1_M (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8-Flash-Next at 30.5 t/s generation and 250.4 t/s prompt processing on 2x RTX 5060 Ti with 32GB system RAM. Setup is llama.cpp with an IQ1_M GGUF (27.58 GB), 98304 context, q8_0 KV cache, tensor split 1,1, and 8 MoE layers offloaded to CPU. User asks whether their settings are correct and what config others run for this model on this hardware.

Oct 5, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 4,096 ctx

reported speed:
29.2 tokens/s generation
quant:
UD-IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.8 27B at 29.18 tok/s decode on an RTX 5060 Ti 16 GB at 4K context. Setup is llama.cpp build ad1de39e0 with UD-IQ3_XXS quant and q4_0 K/V cache, full GPU offload, flash attention, one parallel slot. Baseline decode without speculation is 29.18 tok/s at 4K, 23.62 at 32K and 19.77 at 64K; MTP n=2 gives 56.91/42.53/34.95 and MTP n=3 gives 63.89/45.48/39.06 tok/s at the same contexts. N-gram speculation added only 1.7% (29.66 tok/s). A 96K prompt measured 617.73 prompt tok/s and 36.74 decode tok/s.

Oct 3, 2026
Tone: mixed
reported speed:
35.0 tokens/s generation
quant:
IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports about 35 t/s or more with Qwen3.8 27B GSQ-RCO-Uncensored on an RTX 5060 Ti 16GB. Setup is llama.cpp with IQ3_XXS quant, q4_0 KV cache, 131072 context, and MTP draft plus ngram speculative decoding. The user is on Fedora 44 with an AMD Ryzen 9600x and 16 GB system RAM, and had struggled to get reasonable speed before finding this model.

Oct 1, 2026
reported speed:
47-51 tokens/s generation · 1168.0 tokens/s prompt processing
quant:
UD-IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.5-35B-A3B at 47-51 tok/s generation on an RTX 5060 Ti 16GB over OCuLink in a Proxmox VM. Setup is llama.cpp with UD-IQ3_XXS quant and q4_0 KV cache, fully on the GPU with 348 MiB headroom, at 160K context. Prompt eval of 75K tokens took 64.8 s (1,168 tok/s). The 9B variant reached 40-50 tok/s with UD-Q4_K_XL and about 47-51 tok/s with IQ3_XXS.

Sep 30, 2026
Tone: positive
reported speed:
80.0 tokens/s generation · 4737.5 tokens/s prompt processing
quant:
NVFP4 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Nemotron 3.5 Lightning 30B-A3B at 79.98 t/s generation and 4737.46 t/s prompt processing on 2x RTX 5060 Ti 16GB. Setup is llama.cpp with NVFP4 GGUF weights and q8_0 KV cache, 1048576 token context, flash attention on, no MTP layers. The run processed a 54025-token prompt and generated 2156 tokens; the user notes it fits in 32 GB VRAM with no expert-layer offloading to RAM.

Sep 28, 2026
reported speed:
41.6 tokens/s generation · 1718.3 tokens/s prompt processing
quant:
Q8_0 (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports MiMo 2.6 Distill Qwen 9B at 41.63 t/s generation and 1718.29 t/s prompt processing on an RTX 5060 Ti 16GB. Setup is llama.cpp (Llama UI) with GGUF Q8_0 weights, q4_0 KV cache, 122880-token context, and Flash Attention enabled. The run produced 11979 output tokens over 4 min 47 s from a 1929-token prompt; the first generated version failed and a second review pass by the model produced the final working demo.

Sep 27, 2026

Qwen3.8 27B Swift-Genesis

2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 262,144 ctx

Tone: mixed
reported speed:
76.0 tokens/s generation
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionagenticlong-context

User reports Qwen3.8 27B Swift-Genesis at 76 t/s in their benchmark on 2x RTX 5060 Ti 16GB, with a range of 40-113 t/s and 45-65 t/s under normal agentic work. Setup is llama.cpp with GGUF weights at 262k context, MTP4 speculative decoding, and vision offloaded to system RAM. The user notes the model fits under 17GB with vision, allowing full 262k context on 32GB VRAM, and asks what the tradeoff of this model is compared to other Swift NVFP4 models.

Sep 27, 2026

Occamy 1.0

2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 175,000 ctx

Tone: positive
reported speed:
100.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Occamy 1.0 at around 100 tok/s on two RTX 5060 Ti cards, with enough memory for two independent ~175K context pools for concurrent subagents. Setup uses llama.cpp on the desktop worker box; the cards cost about $400 each. A 96 GB M5 Ultra Mac Studio is planned as the primary inference appliance running Qwen 3.8 Next Flash through oMLX, with DeepSeek V4.1 Flash via API filling in until it arrives.

Sep 24, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 110,000 ctx

reported speed:
5.3 tokens/s generation
quant:
IQ4 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Qwen3.8 27B at ~5.3 t/s on an RTX 5060 Ti 16GB with 16GB single-channel system RAM. Setup is llama.cpp with IQ4 weights, 110k context, Q8 KV cache, and partial GPU/CPU offload; CPU and GPU each sit around 50% utilization. User estimates Q8 weights plus 256k context would need 38-40GB total and drop to roughly 2 t/s, and asks whether quant level, context, or throughput should be prioritized for a local coding-agent worker.

Sep 17, 2026
Tone: mixed
reported speed:
18.3 tokens/s generation · 25.4 tokens/s prompt processing
quant:
Q4 (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next-REAP-320 at 18.3 t/s generation and 25.4 t/s prefill on an RTX 5060 Ti 16GB with 32GB system RAM. Setup is llama.cpp with a Q4 GGUF, 64k context, q4_0 KV cache, --n-cpu-moe 34, --ngl 48, and lazy mmap for the 29.48 GB n-gram embedding. User notes the model is smarter than Qwen3.8 27B but keeps the 27B as daily driver due to 5x faster prompt processing.

Sep 14, 2026
Tone: positive
reported speed:
37.2 tokens/s generation
mtp (multi-token prediction):
off

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User benchmarks multiple models on 16 GB of VRAM. Qwen3.6-35B-A3B-APEX-I-Quality reaches 96.7% HumanEval pass@1, 100% agentic success and 37.2 tok/s. Ornith-1.0-35B IQ4_NL runs at 38 tok/s. Other models tested are Qwen3.8-27B variants, gemma-4-26B-A4B, KAT-Coder-V2.5-Dev and Qwen3.5-9B. Ornith 1.0 35B A3B wins overall.

Sep 9, 2026

Muse 30B Glimmer

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 131,768 ctx

reported speed:
18.0 tokens/s generation
quant:
Q4_K_XL (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a model fits on a single RTX 5060 Ti 16GB at 131k context with a Q4 KV cache. Setup uses GGUF weights only, with no dflash or mmproj loaded. A Q8 KV cache allows about 90k context.

Sep 7, 2026
Tone: positive
reported speed:
40.0 tokens/s generation
quant:
Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

vision

User reports Qwen3.6-35B-A3B MoE at about 40 t/s on an RTX 5060 Ti 16GB with mmproj enabled, and just over 50 t/s without mmproj and with --n-cpu-moe 24. Setup is llama.cpp server in Docker with --cache-type-k/v q8_0, --cache-ram 0, --no-mmap, and --n-cpu-moe 19, or 24 without vision. The user experienced an OOM shutdown before adding --cache-ram 0 and is seeking advice on the setup.

Sep 7, 2026
reported speed:
11.0 tokens/s generation · 200.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports llama.cpp running on 4x RTX 5060 Ti 16GB with DDR4 3200 RAM in 4-channel. Setup uses -ub/-b at 4096.

Sep 7, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 100,000 ctx

Tone: positive
reported speed:
68.3 tokens/s generation
quant:
Q6_K (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at about 68 t/s generation on dual RTX 5060 Ti. Setup is speculative decoding with a Q8 KV cache at 100k context. Vision is enabled but not used.

Sep 7, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 65,536 ctx

Tone: positive
reported speed:
31.3 tokens/s generation · 738.7 tokens/s prompt processing
quant:
IQ4_XS (GGUF)
kv:
Q4
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B on an RTX 5060 Ti 16GB with llama.cpp, comparing the jpetrina IQ4_XS-pure and Unsloth UD-IQ4_XS quants. Short-context generation runs at roughly 45-47 tok/s, dropping to about 31 tok/s after 55K prefill. The Unsloth quant has better fidelity (KLD 0.018 vs 0.0236), while jpetrina wins the multi-turn agent eval (40/80 vs 38/80). Vision works but needs a separate profile. A Q8 baseline was used as a quality control.

Sep 7, 2026
reported speed:
20.0 tokens/s generation · 1000.0 tokens/s prompt processing
quant:
q8 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B q8 GGUF at 20 t/s tg and 1000 pp on a Threadripper rig with 4x RTX 5060 Ti 16GB, VRAM only. User also reports DeepSeek V4 Flash 0731 GGUF at 11 t/s tg and 200 pp with RAM offload. User asks about vLLM tensor parallelism, NVFP4, a CPU upgrade, and other models.

Sep 7, 2026

Qwen3.8 27B

NVIDIA RTX 5060 Ti 16GB · llama.cpp · 200,000 ctx

Tone: mixed
reported speed:
400.0 tokens/s prompt processing
quant:
IQ3_XXS (GGUF)
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports over 200k context on 16 GB of VRAM with Qwen 3.8 27B at the IQ3_XXS quant. Setup is a laptop with Thunderbolt 4 and an Aorus 5060 Ti AI Box eGPU on Windows 11, with the KV cache quantized to q5_1. The user previously ran UD-Q3_K_XL at 140k context. Prompt processing dropped from 700-800 tk/s to 400 tk/s.

Sep 7, 2026

Qwen3.8 27B

3× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 210,000 ctx

Tone: mixed
reported speed:
38-44 tokens/s generation · 1300.0 tokens/s prompt processing
quant:
NVFP4 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B-NVFP4-MTP-GGUF at 210k context on 3x RTX 5060 Ti 16GB in a Dell 5820 workstation with 128GB DDR4 RAM, with prompt processing at 1300-700 t/s and generation at 29-71 t/s, mostly 38-44. Setup is llama.cpp with a Q8 KV cache and vision enabled. The user added a third GPU but reports cooling and stability issues, and is considering a second workstation with subagents for Mixture of Agents. The user also tested Ornith 1.5 35B (MXFP) and Gemma 4 26B NVFP on 16GB VRAM. Use cases are web hosting, office documents, and prompt generation for image and video models.

Sep 7, 2026
Showing 1–20 of 22
Page 1 of 2

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090164NVIDIA RTX 5090111AMD Strix Halo 128GB81NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1382 total
Qwen3.8769
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other211

Records by engine

1062 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other224

Use cases

coding 442agentic 281long-context 202tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding442agentic281long-context202tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active276DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active21Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL48IQ3_XXS38Q436Q8_0354-bit24Q6_K23