llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

GPU: AMD Strix Halo 128GB
Compare setup details (1 active)

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
46.0 tokens/s generation · 1400.0 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8-Flash-Next at 1,400 t/s prefill and 46 t/s decode on a Strix Halo 128GB laptop at 70 W. Setup is Gufo as the engine, with the model running at xhigh effort. The user notes the model ran on an older Halogen version (0.14.0) and Halogen dropped the connection once, so the last part ran on Gufo. The user compares the local model against Claude Opus 5.5 on a coding task, finding the local model's PR better in tests and edge cases, though Opus was 2 to 10 times faster overall.

Oct 6, 2026
reported speed:
50.0 tokens/s generation
quant:
UD-IQ4_XS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 Flash-Next at about 50 tok/s decode with MTP on a Strix Halo 128GB box, even at 116k context. Setup is a forked Gufo engine with the Unsloth UD-IQ4_XS quant (about 89GB) and an MTP sidecar, running in the 120W performance profile. Prefill is 1500+ tok/s from about 3k tokens up, peaking around 1570; on HumanEval prompts decode is about 76 tok/s with MTP versus 28 without. The balanced profile is around 10% less prefill, and very short prompts are slower (about 1350 at 2k) due to a fixed cost per request.

Oct 6, 2026
Tone: positive
reported speed:
34.0 tokens/s generation
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8-27B at 33.96 t/s generation on Strix Halo 128GB. Setup is Gufo with Q4_K_XL GGUF and DFlash2 speculative decoding using a Q4_K_M draft model. The prompt processing figure of 183.3 t/s is noted as low because of the short prompt. User notes it is slower on Q8_0 but still faster than llama.cpp, and is testing the new AUR package.

Oct 5, 2026
Tone: mixed
reported speed:
1124.0 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.8-Flash-Next on a Bosgame M5 (Strix Halo 128GB) across three backends. Gufo v0.7.0 with unsloth UD-Q4_K_XL and MTP reaches 54.0 t/s decode on code, 53.6 t/s on JSON, 29.5 t/s on prose, and 32.9 t/s at 19k context, with prefill 1124/1183/1179 t/s at 4k/19k/38k. strix-llama (ROCm) with ISTA GSQ-RCO IQ3_S and MTP gets 37.9/36.0/24.4/26.6 t/s decode and 734/798/794 t/s prefill; mainline llama.cpp (Vulkan) without MTP gets 27.5/27.5/27.3/25.5 t/s decode and 329/328/285 t/s prefill. GPU memory at 128k context is ~85 GiB for Gufo, ~62 GiB for strix-llama, ~57 GiB for mainline. The user notes Gufo is fastest but only accepts unsloth Q4_K_XL, while IQ3_S on strix-llama leaves room for Gemma 4 26B-A4B alongside. MTP did not work on mainline Vulkan. llama-bench on strix-llama with IQ3_S and no MTP gave pp2048 944 t/s (845 at 32k) and tg128 26.4.

Oct 4, 2026
Tone: positive
reported speed:
40.0 tokens/s generation
quant:
UD_Q4_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports Qwen3.8 Flash-Next at 40 t/s average on agentic tasks on Strix Halo 128GB. Setup is Gufo with UD_Q4_XL quant, VRAM set to 96GB. User notes the speed holds at higher context lengths and recommends setting VRAM to 96GB on Windows.

Oct 3, 2026
Tone: mixed
reported speed:
50.8 tokens/s generation · 2384.0 tokens/s prompt processing
quant:
Q8_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks Qwen3.6-35B-A3B at 50.8 t/s decode and 2384 t/s prefill on a Strix Halo 128GB using the gufo engine with Q8_K_XL quantization. Setup is gufo with Q8_K_XL GGUF and no KV cache quantization, running on the Ryzen AI Max+ 395 with 128 GB unified memory. The user compares gufo against llama.cpp across multiple quantizations and speculative decoding modes. Gufo achieves higher prefill speeds (up to 2702 t/s on Q6_K_XL) but llama.cpp decodes faster on Q6_K_XL by 3% to 9%. With DFlash2 speculative decoding at 7 draft tokens, gufo reaches 78.6 t/s on Q8_K_XL versus llama.cpp's 51.7 t/s. The user notes that speculative output is not bit-identical to plain greedy output on either engine.

Oct 3, 2026
Tone: positive
reported speed:
33.6 tokens/s generation · 1343.0 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextcodingagentic

User benchmarks Qwen3.8 Flash-Next on a ROG Flow Z13 2025 with Strix Halo 128GB, comparing Halogen, Gufo and Rulith at 60W and 93W. Gufo 0.5.0 with UD-Q4_K_XL and MTP Q8_0 draft 7 reaches 1343 tok/s prefill and 33.6 tok/s decode at 200K context at 93W, and 1156 tok/s prefill with 33.0 tok/s decode at 60W. Halogen 0.16.0 is fastest for prefill, hitting 1694 tok/s at 200K and 41.16 tok/s decode at 93W, while Rulith on Windows is fastest at short context with 54.7 tok/s decode at 1K but falls to 32.55 tok/s at 200K. The user notes the comparison is not apples-to-apples because model formats and MTP setups differ between backends.

Oct 3, 2026
reported speed:
32.9 tokens/s generation · 1033.0 tokens/s prompt processing
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codinglong-context

User benchmarks Qwen3.8-Flash-Next on an AMD Strix Halo 128GB (ASUS ROG Flow Z13 GZ302) at 70 W TDP, comparing several engines. The gufo engine with UD-Q4_K_XL weights reaches 32.9 t/s decode and 1,033 t/s prefill at 64k context (32k prompt), with 74% MTP acceptance and 14/14 retrieval. Setup uses gufo (ROCm 7.2.4) with UD-Q4_K_XL GGUF weights and MTP speculative decoding, run via LlamaStash on Arch Linux. Halogen 0.14.0 with native .hgn weights is fastest at 39.3 t/s decode and 1,045 t/s prefill, but is closed source and Docker-only. gufo is open source and loads 4x faster from cold. At 128k context (64k prompt), gufo gets 32.3 t/s decode and 1,047 t/s prefill; at 256k context (130k prompt), 27.9 t/s decode and 997 t/s prefill. In 10 Aider polyglot Python exercises, gufo passes 10/10 in 36.0 min, versus Halogen 21.5 min and CIRU 24.6 min.

Oct 3, 2026
Tone: mixed
reported speed:
70.2 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 70.22 tok/s on a Strix Halo device, reproducing Gufo's published 70.56 tok/s figure. Setup is Gufo 0.4.0 with UD-Q4_K_XL GGUF weights and a DFlash2 Q4_K_M draft model, greedy decoding, thinking off, 128 output tokens, prompt cache off. The 70 tok/s figure comes from a repetitive prompt ("Write the word red exactly 1000 times") where speculative decoding accepts nearly every draft token. On nine ordinary prompts the median is 39.4 tok/s, ranging from 22 to 52. With 8 users the aggregated figure is 122.6 tok/s but wall-clock token delivery is 82 tok/s on the repetitive prompt and 52 on normal prompts. Without the draft model Gufo's docs put it around 12 tok/s. In a head-to-head against halogen 0.13.8 on Qwen3.8 Flash-Next, halogen averaged 43.9 tok/s vs Gufo's 38.2 on normal prompts, and 76.6 vs 63.1 tok/s with 4 users; Gufo was faster at cold prefill (1495 vs 1288 tok/s) and on the repetitive prompt (87.4 vs 56.9 tok/s).

Oct 3, 2026
Tone: mixed
reported speed:
39.4 tokens/s generation
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 39.4 t/s median on normal prompts on a Strix Halo device. Setup is Gufo 0.4.0 with UD-Q4_K_XL GGUF and a DFlash2 Q4_K_M draft model, single user, speculative decoding on. The 70.22 t/s figure comes from a repetitive prompt and is a ceiling; normal prompts range from 22 to 52 t/s. In a head-to-head with halogen 0.13.8 on identical hardware, halogen was about 13% faster for one user and 18% faster with four, while Gufo was 16% faster at prompt processing and much faster on the repetitive prompt.

Oct 1, 2026
Tone: positive
reported speed:
70.6 tokens/s generation · 656.3 tokens/s prompt processing
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B Q4 at up to 70.56 t/s generation and 656.33 t/s prompt processing on a Strix Halo 128GB. Setup is the Gufo inference engine, a new vertically integrated engine built specifically for Strix Halo hardware. At concurrency 8 the model reaches 123 t/s aggregate. The engine also supports DeepSeek V4 Flash, Qwen3.8 Flash-Next, Qwen3 ASR and TTS, Qwen Image 2.1, and MiniMax H3.

Sep 24, 2026
Showing 1–11 of 11
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090114AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409036NVIDIA RTX 5070 Ti30

Records by model

1393 total
Qwen3.8778
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.529
Qwen322
other212

Records by engine

1070 total
llama.cpp569
vLLM153
Strata46
NInfer39
Ollama34
other229

Use cases

coding 450agentic 285long-context 207tool-use 120vision 85summarization 45math 36creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding450agentic285long-context207tool-use120vision85summarization45math36creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B467Qwen3.8 125B · 6B active281DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP495IQ4_XS57Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q437Q8_0354-bit24MXFP423