llamaperf

Qwen3.8 27B

on AMD Strix Halo 128GB · Gufo

Tone: positive
Oct 5, 2026
Throughput
34.0 t/s gen
Quant
Q4_K_XL (GGUF)
System RAM
128 GB

Summary

User reports Qwen3.8-27B at 33.96 t/s generation on Strix Halo 128GB. Setup is Gufo with Q4_K_XL GGUF and DFlash2 speculative decoding using a Q4_K_M draft model. The prompt processing figure of 183.3 t/s is noted as low because of the short prompt. User notes it is slower on Q8_0 but still faster than llama.cpp, and is testing the new AUR package.