llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: AMD Radeon 890M
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
14.8 tokens/s generation · 212.1 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.5-35B-A3B at 14.85 t/s generation and 212.05 t/s prompt processing on an AMD Radeon 890M (gfx1150) iGPU. Setup is llama.cpp (build a0ed91a44) with a Q4_K_XL GGUF, 99 layers offloaded, comparing Vulkan and ROCm backends with flash attention on and off. ROCm reaches 227.54 t/s prefill and 12.68 t/s decode without flash attention, and 229.81 t/s prefill with 13.18 t/s decode with it. A second build with unified memory and ROCm flash attention gave similar results. The same machine also ran llama-2-7b Q4_0 at 17.14 t/s decode on Vulkan and 15.52 t/s on ROCm.

Oct 9, 2026
Showing 1–1 of 1
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090176NVIDIA RTX 5090125AMD Strix Halo 128GB102NVIDIA DGX Spark71NVIDIA RTX Pro 6000 Blackwell57NVIDIA RTX 5060 Ti 16GB56AMD Radeon AI PRO R9700 32GB50NVIDIA RTX 3060 12GB49NVIDIA RTX 409039NVIDIA RTX 5070 Ti33

Records by model

1576 total
Qwen3.8872
Qwen3.6182
DeepSeek V4 Flash134
Gemma 470
Qwen3.543
DeepSeek V4.1 Flash27
other248

Records by engine

1209 total
llama.cpp619
vLLM179
Strata58
NInfer44
Ollama37
other272

Use cases

coding 505agentic 329long-context 243tool-use 137vision 97summarization 48math 40creative-writing 34multilingual 21text-generation 9rp 7reasoning 3
coding505agentic329long-context243tool-use137vision97summarization48math40creative-writing34

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509082RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX48RTX 309036RTX 5090 Laptop 24GB34V100 32GB33RTX 409032Radeon AI PRO R9700 32GB30RX 7800 XT 16GB30

Reports by model size

Qwen3.8 27B512Qwen3.8 125B · 6B active328DeepSeek V4 Flash 284B · 13B active126Qwen3.6 35B · 3B active110Qwen3.6 27B70Gemma 4 26B · 4B active28DeepSeek V4.1 Flash 552B · 16B active27GLM-5.3 320B · 18B active21

Quants

Q4_K_M132NVFP4112Q4_K_XL63IQ4_XS62UD-Q4_K_XL58Q442IQ3_XXS40Q8_036MXFP4334-bit30