llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: AMD RX 9070 XT 16GB
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: mixed
reported speed:
62.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentictool-usesummarization

User reports Qwen3.6 35B A3B at 62 t/s on an AMD RX 9070 XT 16GB, with vLLM on ROCm reaching 48 t/s in the same comparison. Setup is llama.cpp with the Vulkan backend, 32k context, on a Ryzen 7 9800X3D with 64GB DDR5; the MoE model spills part of its weights to system RAM. User notes the setup is early and numbers are not settled, that 16GB VRAM rules out big dense models, and that the local model needs supervision.

Oct 5, 2026

Qwen3.8 27B

AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx

Tone: positive
reported speed:
42.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen 3.8 27B at 42 t/s on an AMD RX 9070 XT 16GB. Setup is llama.cpp with MTP2 speculative decoding and a 64K context, using about 1 GB of working RAM during evaluation. MTP2 delivered +41% to +53% throughput depending on the quantization, but the user notes the fastest configuration was not necessarily best for reasoning quality.

Sep 29, 2026

Qwen3.8 27B

AMD RX 9070 XT 16GB · llama.cpp · 131,072 ctx

reported speed:
50.0 tokens/s generation
quant:
Q6 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 50 t/s for Qwen 3.8 27B on an RX 9070 XT 16GB. Setup is llama.cpp (Vulkan build) with a Q6 quant, 131072 context, Q8 K cache and Q4 V cache, and DFlash2 speculative decoding with 7 max drafts. The rig also has an R9700 AI Pro 32GB and 64GB DDR5 6000; the user asks what speeds and flags others get, especially on Windows/AMD.

Sep 28, 2026

Qwen3.8 27B Swift-1.5

AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx

Tone: positive
reported speed:
43.7 tokens/s generation · 700.8 tokens/s prompt processing
quant:
IQ3_S (GGUF)
kv:
q5_0-q4_1

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

tool-uselong-context

User reports Qwen3.8 27B (Swift-1.5-GSQ-RCO IQ3_S with MTP) at 43.7 tok/s decode and 700.8 t/s prefill on an RX 9070 XT 16GB. Setup is llama.cpp (server-vulkan b11176) with IQ3_S weights and a q5_0-q4_1 KV cache at 65536 context, one card. Decode is the mean of 3 runs of a fixed 512-token generation; prefill is one cold 30,461-token prompt. At 131072 context the same Vulkan launch gives 37.4 tok/s decode and 701.6 t/s prefill, while the pinned ROCm image (b10884) cannot load at 128K and decodes 33.5 tok/s at 65536. The user attributes the ROCm failure to a missing FlashAttention vector kernel for the q5_0/q4_1 K/V pair, which promotes both to f16 and exhausts the 16GB card.

Sep 27, 2026

Qwen3.8 27B

2× AMD RX 9070 XT 16GB · llama.cpp · 98,304 ctx

reported speed:
24.7 tokens/s generation · 2061.1 tokens/s prompt processing
quant:
Q4_K_M (GGUF)
kv:
f8_e4m3

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentictool-usevisionlong-context

User reports Qwen3.8-27B at 2061.11 t/s prompt and 24.70 t/s decode on two Radeon RX 9070 XT 16 GB cards under ROCm 7.2 on Windows. Setup is a llama.cpp fork with Q4_K_M GGUF, f8_e4m3 KV cache, FlashAttention, batch 8192 / ubatch 1024, layer split across both GPUs, one server slot, and cold prompt processing at 98,304 context (L3 lane). With MTP n3 the same L1 lane gives 1817.81 t/s prompt and 36.73 t/s decode at 46.8% acceptance. Vulkan rows are lower on prompt processing (L1 1569.17, L2 1470.80, L3 1243.99 t/s) and lead only the L1 decode rows. Linux ROCm 10 results reach 2253.06 t/s prompt with MXFP4-requant and 50.13 t/s decode with Q4_K_M plus MTP n3.

Sep 27, 2026
reported speed:
706.0 tokens/s prompt processing
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B at 706.0 t/s prompt processing on an AMD Radeon RX 9070 XT 16GB under ROCm. Setup is llama.cpp with Q4_K_M GGUF, flash attention enabled, -ncmoe 40 and -ub 512, on a Ryzen 9 5950X with 64 GB DDR4-3200. The same configuration under Vulkan reached 333.4 t/s prompt and 29.3 t/s generation; the user retired Vulkan in favour of ROCm.

Sep 27, 2026

Qwen3.5 9B

AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx

Tone: positive
reported speed:
62.0 tokens/s generation
quant:
Q6_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.5-9B at 62 t/s on an AMD Radeon RX 9070 XT (RDNA4/gfx1201) using llama-server with the Vulkan backend. Setup is llama.cpp with GGUF Q6_K_XL weights at 65536 tokens of context, single request, no speculation. The same model under vLLM ROCm 7.2 with FP8 weights reached 48 t/s, which the user attributes to vLLM lacking native gfx1201 kernel support and falling back to FP32 dequantization. The user reports llama-server Vulkan is 29% faster on this hardware and recommends it over vLLM until RDNA4 support lands upstream.

Sep 27, 2026

Qwopus3.5 9B Coder

AMD RX 9070 XT 16GB · 128,000 ctx

reported speed:
150.0 tokens/s generation
quant:
Q8_0 (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User compares Qwopus3.5-9B-Coder at Q8_0 against Qwen3.8-27B at IQ4_XS. The 9B runs at 150+ t/s with 128K context and a BF16 KV cache. The 27B is limited to ~80K context with Kvarn4 and MTP with a Q4 cache at 20 t/s. User asks whether the 9B is still worth using for coding in 2026.

Sep 23, 2026

Qwen3.8 27B Swift

2× AMD RX 9070 XT 16GB · llama.cpp · 131,072 ctx

Tone: mixed
reported speed:
50.0 tokens/s generation
quant:
Q6 (GGUF)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen 3.8 Swift at a ceiling of 50 t/s on an RX 9070 XT 16GB and a Radeon AI PRO R9700 32GB. Setup is llama.cpp with the Vulkan backend, a Q6 quant, 131K context, and Q8 KV cache for both K and V, with a DFlash2 drafter. User asks whether others get more than 50 t/s on AMD 9070 XT or R9700 hardware with a usable context size, and whether a different OS could add 30+ t/s.

Sep 21, 2026
Showing 1–9 of 9
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti30

Records by model

1397 total
Qwen3.8781
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other213

Records by engine

1074 total
llama.cpp570
vLLM153
Strata47
NInfer40
Ollama34
other230

Use cases

coding 453agentic 287long-context 208tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding453agentic287long-context208tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B468Qwen3.8 125B · 6B active283DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22GLM-5.3 320B · 18B active18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24Q6_K23