llamaperf
Oct 3, 2026
Throughput
32.9 t/s gen · 1033.0 t/s pp
Quant
UD-Q4_K_XL (GGUF)
System RAM
128 GB

Use cases

codinglong-context

Summary

User benchmarks Qwen3.8-Flash-Next on an AMD Strix Halo 128GB (ASUS ROG Flow Z13 GZ302) at 70 W TDP, comparing several engines. The gufo engine with UD-Q4_K_XL weights reaches 32.9 t/s decode and 1,033 t/s prefill at 64k context (32k prompt), with 74% MTP acceptance and 14/14 retrieval. Setup uses gufo (ROCm 7.2.4) with UD-Q4_K_XL GGUF weights and MTP speculative decoding, run via LlamaStash on Arch Linux. Halogen 0.14.0 with native .hgn weights is fastest at 39.3 t/s decode and 1,045 t/s prefill, but is closed source and Docker-only. gufo is open source and loads 4x faster from cold. At 128k context (64k prompt), gufo gets 32.3 t/s decode and 1,047 t/s prefill; at 256k context (130k prompt), 27.9 t/s decode and 997 t/s prefill. In 10 Aider polyglot Python exercises, gufo passes 10/10 in 36.0 min, versus Halogen 21.5 min and CIRU 24.6 min.