Qwen3.8 27B
on AMD Strix Halo 128GB · Gufo
Sep 24, 2026
Summary
User reports Qwen3.8 27B Q4 at up to 70.56 t/s generation and 656.33 t/s prompt processing on a Strix Halo 128GB.
Setup is the Gufo inference engine, a new vertically integrated engine built specifically for Strix Halo hardware.
At concurrency 8 the model reaches 123 t/s aggregate. The engine also supports DeepSeek V4 Flash, Qwen3.8 Flash-Next, Qwen3 ASR and TTS, Qwen Image 2.1, and MiniMax H3.