llamaperf

Qwen3.8 27B

on AMD Strix Halo 128GB · Gufo

Tone: mixed
Oct 3, 2026
Throughput
70.2 t/s gen
Quant
UD-Q4_K_XL (GGUF)

Summary

User reports Qwen3.8 27B at 70.22 tok/s on a Strix Halo device, reproducing Gufo's published 70.56 tok/s figure. Setup is Gufo 0.4.0 with UD-Q4_K_XL GGUF weights and a DFlash2 Q4_K_M draft model, greedy decoding, thinking off, 128 output tokens, prompt cache off. The 70 tok/s figure comes from a repetitive prompt ("Write the word red exactly 1000 times") where speculative decoding accepts nearly every draft token. On nine ordinary prompts the median is 39.4 tok/s, ranging from 22 to 52. With 8 users the aggregated figure is 122.6 tok/s but wall-clock token delivery is 82 tok/s on the repetitive prompt and 52 on normal prompts. Without the draft model Gufo's docs put it around 12 tok/s. In a head-to-head against halogen 0.13.8 on Qwen3.8 Flash-Next, halogen averaged 43.9 tok/s vs Gufo's 38.2 on normal prompts, and 76.6 vs 63.1 tok/s with 4 users; Gufo was faster at cold prefill (1495 vs 1288 tok/s) and on the repetitive prompt (87.4 vs 56.9 tok/s).