llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: AMD RX 9060 XT 16GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

reported speed:
66.0 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B at 66.04 tokens/s on an AMD Radeon RX 9060 XT 16 GB with 32 GB of system RAM. Setup is llama.cpp (Vulkan backend) with UD-Q4_K_XL GGUF weights, q8_0 KV cache, 40k context, and MTP speculative decoding drafting up to 3 tokens per step. The 35B MoE model does not fit in 16 GB of VRAM, so the expert weights of the first 20 layers run on the CPU. Run on a Ryzen 7 5800X3D desktop under SteamOS, single slot, with flash attention and prefix cache reuse enabled.

Oct 6, 2026
reported speed:
30.0 tokens/s generation
quant:
IQ3_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User asks whether Qwen3.8 Flash-Next can run on an RX 9060 XT 16GB with 32GB DDR5-6000. They currently run Swift-1.5-Qwen3.8 at IQ3_S at 30 t/s on the same machine.

Oct 3, 2026
reported speed:
50.0 tokens/s generation · 850.0 tokens/s prompt processing
quant:
IQ3_S (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 850 t/s prefill and 50 t/s generation with MTP at temperature 0 on an AMD RX 9060 XT 16GB, running Qwen3.8-27B-GSQ-RCO IQ3_S in a custom inference engine. MTP is used for the generation figure; the user notes MTP with temperature above 0 is not developed yet. User asks for ideas to systematically test the engine, having already run KL divergence against BF12 on CPU, bit correctness checks, needle-in-a-haystack at 25/50/75/90% key positions with distractors (also against llama.cpp), and HumanEval (92/93 passed so far).

Oct 3, 2026
reported speed:
33.5 tokens/s generation · 1210.3 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen 2.5-Coder 14B Instruct Q4_K_M at 33.51 t/s generation and 1210.29 t/s prompt processing on an AMD RX 9060 XT 16GB via Vulkan. Setup is llama.cpp build c4ae9a88f8 with -ngl 99 and -fa 1; the same model on ROCm reaches 30.79 t/s generation and 1211.32 t/s prompt processing. A Qwen 3.5 9B Q4_K_M run reaches 50.15 t/s generation and 1904.16 t/s prompt processing on Vulkan, and a 14B quant sweep gives about 32, 28 and 20 tok/s for Q4_K_M, Q5_K_M and Q8_0.

Sep 29, 2026

Ornith1.5 9B

AMD RX 9060 XT 16GB · llama.cpp · 262,144 ctx

Tone: positive
reported speed:
36.0 tokens/s generation · 950.0 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Ornith-1.5-9B at about 36 t/s generation and about 950 t/s prompt eval on an AMD RX 9060 XT 16GB with llama.cpp. Throughput drops to about 25 t/s generation and about 500 t/s prompt eval during long tasks. The run lasted about 3.5 hours on a coding agent task. The user contrasts this with Qwen 3.8 27B, which was frustratingly slow on the same hardware.

Sep 7, 2026
Showing 1–5 of 5
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423