llamaperf
Oct 9, 2026
Throughput
14.8 t/s gen · 212.1 t/s pp
Quant
Q4_K_XL (GGUF)

Summary

User benchmarks Qwen3.5-35B-A3B at 14.85 t/s generation and 212.05 t/s prompt processing on an AMD Radeon 890M (gfx1150) iGPU. Setup is llama.cpp (build a0ed91a44) with a Q4_K_XL GGUF, 99 layers offloaded, comparing Vulkan and ROCm backends with flash attention on and off. ROCm reaches 227.54 t/s prefill and 12.68 t/s decode without flash attention, and 229.81 t/s prefill with 13.18 t/s decode with it. A second build with unified memory and ROCm flash attention gave similar results. The same machine also ran llama-2-7b Q4_0 at 17.14 t/s decode on Vulkan and 15.52 t/s on ROCm.