llamaperf

Qwen3.8 27B

on AMD Strix Halo 128GB · Gufo

Tone: mixed
Oct 1, 2026
Throughput
39.4 t/s gen
Quant
UD-Q4_K_XL (GGUF)

Summary

User reports Qwen3.8 27B at 39.4 t/s median on normal prompts on a Strix Halo device. Setup is Gufo 0.4.0 with UD-Q4_K_XL GGUF and a DFlash2 Q4_K_M draft model, single user, speculative decoding on. The 70.22 t/s figure comes from a repetitive prompt and is a ceiling; normal prompts range from 22 to 52 t/s. In a head-to-head with halogen 0.13.8 on identical hardware, halogen was about 13% faster for one user and 18% faster with four, while Gufo was 16% faster at prompt processing and much faster on the repetitive prompt.