Qwen3.8 125B (6B active) Flash-Next
on AMD Strix Halo 128GB · Strata · 8,192 ctx
Oct 6, 2026
Summary
User reports Qwen3.8-Flash-Next at 53.8 t/s output and 1293 t/s prompt processing on Strix Halo 128GB.
Setup is Strata with UD-IQ4_XS weights at 8K context; the release adds official Strix Halo support, marked experimental and Linux only.
At 64K the run gives 46.6 t/s output and 1370 t/s prompt processing, and at 128K 51.4 t/s output and 1320 t/s prompt processing. The user says it can go up to 1M context length without big speed loss.