llamaperf

How open-weight LLMs run on your hardware

75 performance reports, crowdsourced from the community.

VRAM Calculator

Pick your GPU - see which models fit, at which quant, and how fast they run.

Open calculator →

Gemma 4

M5 32GB · MLX · 130,173 ctx

throughput:
3029.0 t/s pp

User built custom w8a8 kernels for M5 MacBook Air. Baseline prefill 2193 tps, improved to 3029 tps for 130k tokens. Also mentions llama.cpp for Macs.

Qwen3.6 27B

RTX 5060 Ti 16GB · 256,000 ctx

Tone: positive
throughput:
52.2 t/s gen · 608.0 t/s pp
quant:
Q8
kv:
F16
mtp (multi-token prediction):
on
coding

Benchmark on Vast AI instance with 4x RTX 5060 Ti 16GB. Q8 quant, FP16 KV cache, MTP enabled. 256K context. Cold prefill 608 t/s, decode 52.2 t/s. User considers this excellent for $2K hardware.

Kimi K2.6

RTX 5090 · llama.cpp

throughput:
471.4 t/s pp
quant:
IQ3_M (gguf)

LLM prompt processing benchmark with Kimi K2.5 IQ3_M (80GB offload) at 500W. RTX 5090 achieved 471.40 t/s PP. Also tested GLM 5.1 IQ4_NL (70GB offload) at 574.98 t/s PP. Comparison with RTX 6000 PRO MaxQ shunt modded.

Qwen2.5 27B

RTX 3090 · llama.cpp

Tone: positive
throughput:
70.0 t/s gen · 1850.0 t/s pp
quant:
Q6_K_XL (gguf)
coding

Multi-token prediction enabled. 96GB total VRAM (24+24? but user says 96GB system). Reliable for code generation and codebase ingestion.

Qwen3.6 27B

RTX 3090 Ti · llama.cpp · 196,608 ctx

Tone: positive
throughput:
100.0 t/s gen
quant:
Q8_0 (gguf)

Tensor split-mode improved t/s from 70+ to 100+. Peak 130 t/s reported. Power draw 750W+.

Tone: positive
quant:
Q4_K_M (gguf)
vision

Champion model in vision benchmark. Best quality and stability with thinking disabled. 90/90 successful runs. Speed: 70 s/img. Tested on Apple M2 Max 96GB with llama.cpp b9690.

Qwen3.6 27B

RTX 5060 Ti 16GB · llama.cpp · 131,072 ctx

Tone: positive
throughput:
19.0 t/s gen
quant:
IQ4_XS (gguf)
kv:
F16

User reports that offloading KV cache to RAM (with -nkvo) allows fitting the whole model on GPU with f16 KV cache, achieving 19 tps peak and 14 tps during long generation at 65k context. With 128k context and 63 layers on GPU, speed remained similar. KV cache quant to RAM didn't improve performance.

Gemma 4

RX 7900 XTX · llama-swap

Tone: positive

Benchmark of Gemma 4 QAT vs regular quants on AMD 7900 XTX. No token/s reported, but wall clock times show significant speedups (e.g., 12B QAT 45% faster, 83% throughput increase). Quality reported identical. Models tested: 12B, 26B, 31B, E4B.

Benchmark of abliteration tools (Apostate, Huihui, Heretic) on Qwen 2.5 7B. Evaluated with lm-evaluation-harness via vLLM 0.19.0, bf16 on RTX 5090 32GB. Reports MMLU, GSM8K, HellaSwag, ARC Challenge, WinoGrande, TruthfulQA MC2, PiQA, LAMBADA ppl, HarmBench ASR, KL divergence. No tokens/sec reported.

Tone: positive
throughput:
138.0 t/s gen
coding

Benchmark of Gemma 4 26B-A4B vs 12B on RTX 4090. 26B-A4B used 15GB VRAM, 138 tok/s; 12B used 9GB, 80 tok/s. 26B-A4B won every scene and ran ~1.7x faster. 12B ideal for 16GB laptop.

Qwen3.6 27B

RTX 3090 · Ollama · 32,000 ctx

Tone: mixed
quant:
Q6_K (gguf)
codingagentic

User replaced Claude with Qwen3.6-27B in multi-agent orchestrator for 2 weeks. Plan generation good, tool-call reliability poor (12% format error), long-context drift past ~14k tokens, cascade-failure handling weak. Viable as reasoning layer but not execution layer.

quant:
4bit
agenticcoding

User currently runs Qwen3.6-35B-A3B-4bit on M3 Max 128GB for production sub-agent delegations. Also mentions GLM-5.1 for orchestration. Considering building a 5090 rig.

Qwen3.6 27B

RTX 3090 · 128,000 ctx

throughput:
104.0 t/s gen · 1399.0 t/s pp
quant:
Q8
kv:
F16

User recommends getting enough GPUs to avoid VRAM hacks. Uses 2x RTX 3090s.

Tone: positive
throughput:
70.0 t/s gen
creative-writing

User mentions using Qwen3.6-27B on dual RTX 3090s for generating interactive HTML content inline with chat. Reports ~70 t/s.

Qwen3.6 27B

RTX 3060 12GB · llama.cpp · 64,000 ctx

Tone: positive
throughput:
43.3 t/s gen · 456.1 t/s pp
quant:
Q4_K_S (gguf)

Dual RTX 3060 setup with tensor parallel. MTP enabled. Context 64k. Prefill 456 t/s, generation 43.26 t/s at 12k context. Without MTP, context 96k, generation 31 t/s. User praises value and stability of CUDA.

throughput:
243.9 t/s gen · 13809.2 t/s pp
quant:
Q4_K_M (gguf)
kv:
Q8

Benchmark of FWHT CUDA implementation for kv-cache quantization. Results show 1-2% pp boost and 7-9% tg boost on Gemma 4 26B.A4B Q4_K_M with -ctk q8_0 -ctv q8_0. pp2048 and tg128 values reported; highest t/s from cuda-fwt column.

throughput:
3500.0 t/s gen · 30000.0 t/s pp

Two benchmarks: Qwen3.6 27B BF16 and Qwen3.6 35B BF16. For 35B, best gen tps 3500 at 128 concurrency with MTP off, prompt tps 30000. Also tested 27B with MTP on/off.

Tone: positive
visionsummarization

Model based on Qwen3.5-4B. Trained on 8xH100 for 3 days. Supports Safetensors, GGUF, MLX weights. Requires as little as 4GB VRAM. Multiple quantizations available (GPTQ, W8A8, FP8, Q4, Q6). Tested with vLLM, SGLang, llama.cpp.

Showing 120 of 90
Page 1 of 5

Community benchmarks snapshot

90 records · 24 GPUs · 8 model families · 5 engines

Records by GPU

RTX 5090 12 RTX 3090 11 RTX 3060 12GB 5 M5 Max 128GB 5 RX 7900 XTX 4 AMD Strix Halo 128GB 4 RTX 5060 Ti 16GB 4 RTX 4090 4 RTX Pro 6000 Blackwell 3 H100 80GB 3

Records by model

84 total
Qwen3.656
Gemma 419
Qwen2.53
Kimi K2.62
GLM-5.11
Qwen3.51
other2

Records by engine

61 total
llama.cpp32
vLLM13
Ollama9
MLX5
LM Studio2

Use cases

coding 36agentic 11text-generation 9tool-use 7summarization 6long-context 5vision 4creative-writing 3math 3multilingual 1
coding 36 agentic 11 text-generation 9 tool-use 7 summarization 6 long-context 5 vision 4 creative-writing 3

Avg gen t/s by GPU

Pro 6000 3500 5090 648 4070 Ti Super 110 3090 Ti 100 4090 98 H100 80GB 85 M5 Max 64GB 64 3090 61 5080 56 4070 55

Avg gen t/s by model

Qwen3.6 199 Gemma 4 102 Qwen2.5 49 Kimi K2.6 10 Llama 3.1 1

Quants

Q4_K_M 16 IQ4_XS 6 Q8 5 NVFP4 4 Q4 3 Q8_0 3 Q6_K 2 Q5_K_S 2 UD-Q5_K_XL 1 float8 1