llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: AMD Radeon 780M
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
25.0 tokens/s generation · 208.7 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma4 26B Q4_K_M at ~25.00 t/s generation and ~208.72 t/s prompt processing on an AMD Radeon 780M iGPU. Setup is llama.cpp with Vulkan backend, Q4_K_M GGUF, -ngl 99, on a MINISFORUM UM890 Pro mini PC running Ubuntu 24.04. A CLI smoke test gave ~23.4 t/s generation and ~37.3 t/s prompt; a no-reasoning run gave ~24.4 t/s generation and ~117.6 t/s prompt. Ollama on the same box was around 4.5 t/s generation, roughly a 5x-6x uplift. Ollama's installed Gemma4 blob could not be loaded directly in upstream llama.cpp due to a tensor count mismatch.

Sep 26, 2026
Tone: mixed
reported speed:
3.2 tokens/s generation · 32.5 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at 3.20 t/s generation and 32.45 t/s prompt processing on a Ryzen 9 7940HS with Radeon 780M iGPU and 24GB RAM. Setup is Ollama on two machines: a Minisforum UM790 Pro and a home build with the same Ryzen 9 7940HS, both with 24GB RAM. The Minisforum run reached 4.73 t/s generation and 20.52 t/s prompt processing. Both runs stopped mid-task and follow-up prompts returned no output. User expected low throughput but not the failures.

Sep 18, 2026
Tone: positive
reported speed:
21.1 tokens/s generation · 287.3 tokens/s prompt processing
quant:
Q8_0
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6 35B-A3B Q8_0 on a Ryzen 7 260 with a 780M iGPU and 64 GB DDR5, reaching pp8192 287.33 t/s and tg128 21.06 t/s. Setup is the Vulkan backend. Gemma 4 31B Q8_0 on the same hardware gives pp8192 51.59 t/s and tg128 2.46 t/s. With MTP, Gemma 4 31B reaches tg ~5.76 t/s, and Qwen3.6 35B-A3B with MTP and partial offloading reaches tg ~34.85 t/s. The user mentions an RTX 5060 8GB as a bonus for MoE partial offloading.

Sep 7, 2026
reported speed:
18.4 tokens/s generation · 311.4 tokens/s prompt processing
quant:
Q8 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks ROCm against Vulkan on a Radeon 780m iGPU. ROCm gives a 50% pp speedup for dense models. The user also reports Qwen3.8 27B results.

Sep 7, 2026
Tone: mixed
reported speed:
100.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports ROCm 7.14 on a Radeon 780M with llama.cpp is fast but unstable and crashes. Setting AMD_SERIALIZE_KERNEL=3 stabilizes it but drops prompt processing to about 100 t/s, down from 200-300 t/s without the workaround. User asks for advice.

Sep 7, 2026
Showing 1–5 of 5
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti30

Records by model

1397 total
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213

Records by engine

1074 total
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230

Use cases

coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding453agentic287long-context208tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B468Qwen3.8 125B · 6B active283DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22GLM-5.3 320B · 18B active18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24Q6_K23