llamaperf

Qwen3.8 125B (6B active) Flash-Next

on AMD Strix Halo 128GB · Strata · 8,192 ctx

Tone: positive
Oct 6, 2026
Throughput
53.8 t/s gen · 1293.0 t/s pp
Quant
UD-IQ4_XS (GGUF)
System RAM
128 GB

Summary

User reports Qwen3.8-Flash-Next at 53.8 t/s output and 1293 t/s prompt processing on Strix Halo 128GB. Setup is Strata with UD-IQ4_XS weights at 8K context; the release adds official Strix Halo support, marked experimental and Linux only. At 64K the run gives 46.6 t/s output and 1370 t/s prompt processing, and at 128K 51.4 t/s output and 1320 t/s prompt processing. The user says it can go up to 1M context length without big speed loss.