llamaperf
Oct 6, 2026
Throughput
50.0 t/s gen
Quant
UD-IQ4_XS (GGUF)
System RAM
128 GB

Summary

User reports Qwen3.8 Flash-Next at about 50 tok/s decode with MTP on a Strix Halo 128GB box, even at 116k context. Setup is a forked Gufo engine with the Unsloth UD-IQ4_XS quant (about 89GB) and an MTP sidecar, running in the 120W performance profile. Prefill is 1500+ tok/s from about 3k tokens up, peaking around 1570; on HumanEval prompts decode is about 76 tok/s with MTP versus 28 without. The balanced profile is around 10% less prefill, and very short prompts are slower (about 1350 at 2k) due to a fixed cost per request.