Oct 4, 2026
Summary
User benchmarks Qwen3.8-Flash-Next on a Bosgame M5 (Strix Halo 128GB) across three backends. Gufo v0.7.0 with unsloth UD-Q4_K_XL and MTP reaches 54.0 t/s decode on code, 53.6 t/s on JSON, 29.5 t/s on prose, and 32.9 t/s at 19k context, with prefill 1124/1183/1179 t/s at 4k/19k/38k. strix-llama (ROCm) with ISTA GSQ-RCO IQ3_S and MTP gets 37.9/36.0/24.4/26.6 t/s decode and 734/798/794 t/s prefill; mainline llama.cpp (Vulkan) without MTP gets 27.5/27.5/27.3/25.5 t/s decode and 329/328/285 t/s prefill. GPU memory at 128k context is ~85 GiB for Gufo, ~62 GiB for strix-llama, ~57 GiB for mainline. The user notes Gufo is fastest but only accepts unsloth Q4_K_XL, while IQ3_S on strix-llama leaves room for Gemma 4 26B-A4B alongside. MTP did not work on mainline Vulkan. llama-bench on strix-llama with IQ3_S and no MTP gave pp2048 944 t/s (845 at 32k) and tg128 26.4.