Qwen3.5 35B (3B active)
on AMD Radeon 890M · llama.cpp
Oct 9, 2026
Summary
User benchmarks Qwen3.5-35B-A3B at 14.85 t/s generation and 212.05 t/s prompt processing on an AMD Radeon 890M (gfx1150) iGPU.
Setup is llama.cpp (build a0ed91a44) with a Q4_K_XL GGUF, 99 layers offloaded, comparing Vulkan and ROCm backends with flash attention on and off.
ROCm reaches 227.54 t/s prefill and 12.68 t/s decode without flash attention, and 229.81 t/s prefill with 13.18 t/s decode with it. A second build with unified memory and ROCm flash attention gave similar results. The same machine also ran llama-2-7b Q4_0 at 17.14 t/s decode on Vulkan and 15.52 t/s on ROCm.