llamaperf

Qwen3.8 27B

on AMD Strix Halo 128GB · Gufo

Tone: positive
Sep 24, 2026
Throughput
70.6 t/s gen · 656.3 t/s pp
Quant
Q4
System RAM
128 GB

Summary

User reports Qwen3.8 27B Q4 at up to 70.56 t/s generation and 656.33 t/s prompt processing on a Strix Halo 128GB. Setup is the Gufo inference engine, a new vertically integrated engine built specifically for Strix Halo hardware. At concurrency 8 the model reaches 123 t/s aggregate. The engine also supports DeepSeek V4 Flash, Qwen3.8 Flash-Next, Qwen3 ASR and TTS, Qwen Image 2.1, and MiniMax H3.