llamaperf

Qwen3.8 125B (6B active) Flash-Next

on AMD Strix Halo 128GB · Gufo · 200,000 ctx

Tone: positive
Oct 3, 2026
Throughput
33.6 t/s gen · 1343.0 t/s pp
Quant
UD-Q4_K_XL (GGUF)
System RAM
128 GB

Use cases

long-contextcodingagentic

Summary

User benchmarks Qwen3.8 Flash-Next on a ROG Flow Z13 2025 with Strix Halo 128GB, comparing Halogen, Gufo and Rulith at 60W and 93W. Gufo 0.5.0 with UD-Q4_K_XL and MTP Q8_0 draft 7 reaches 1343 tok/s prefill and 33.6 tok/s decode at 200K context at 93W, and 1156 tok/s prefill with 33.0 tok/s decode at 60W. Halogen 0.16.0 is fastest for prefill, hitting 1694 tok/s at 200K and 41.16 tok/s decode at 93W, while Rulith on Windows is fastest at short context with 54.7 tok/s decode at 1K but falls to 32.55 tok/s at 200K. The user notes the comparison is not apples-to-apples because model formats and MTP setups differ between backends.