Qwen3.8 125B (6B active) Flash-Next
on AMD Strix Halo 128GB · Gufo · 65,536 ctx
Oct 3, 2026
Use cases
codinglong-context
Summary
User benchmarks Qwen3.8-Flash-Next on an AMD Strix Halo 128GB (ASUS ROG Flow Z13 GZ302) at 70 W TDP, comparing several engines. The gufo engine with UD-Q4_K_XL weights reaches 32.9 t/s decode and 1,033 t/s prefill at 64k context (32k prompt), with 74% MTP acceptance and 14/14 retrieval.
Setup uses gufo (ROCm 7.2.4) with UD-Q4_K_XL GGUF weights and MTP speculative decoding, run via LlamaStash on Arch Linux.
Halogen 0.14.0 with native .hgn weights is fastest at 39.3 t/s decode and 1,045 t/s prefill, but is closed source and Docker-only. gufo is open source and loads 4x faster from cold. At 128k context (64k prompt), gufo gets 32.3 t/s decode and 1,047 t/s prefill; at 256k context (130k prompt), 27.9 t/s decode and 997 t/s prefill. In 10 Aider polyglot Python exercises, gufo passes 10/10 in 36.0 min, versus Halogen 21.5 min and CIRU 24.6 min.