Qwen3.8 125B (6B active) Flash-Next
on AMD Strix Halo 128GB · Gufo · 200,000 ctx
Oct 3, 2026
Use cases
long-contextcodingagentic
Summary
User benchmarks Qwen3.8 Flash-Next on a ROG Flow Z13 2025 with Strix Halo 128GB, comparing Halogen, Gufo and Rulith at 60W and 93W.
Gufo 0.5.0 with UD-Q4_K_XL and MTP Q8_0 draft 7 reaches 1343 tok/s prefill and 33.6 tok/s decode at 200K context at 93W, and 1156 tok/s prefill with 33.0 tok/s decode at 60W.
Halogen 0.16.0 is fastest for prefill, hitting 1694 tok/s at 200K and 41.16 tok/s decode at 93W, while Rulith on Windows is fastest at short context with 54.7 tok/s decode at 1K but falls to 32.55 tok/s at 200K. The user notes the comparison is not apples-to-apples because model formats and MTP setups differ between backends.