Qwen3.8 27B
on AMD Strix Halo 128GB · Gufo
Oct 5, 2026
Summary
User reports Qwen3.8-27B at 33.96 t/s generation on Strix Halo 128GB.
Setup is Gufo with Q4_K_XL GGUF and DFlash2 speculative decoding using a Q4_K_M draft model. The prompt processing figure of 183.3 t/s is noted as low because of the short prompt.
User notes it is slower on Q8_0 but still faster than llama.cpp, and is testing the new AUR package.