Qwen3.6 35B (3B active)
on AMD RX 9070 XT 16GB · llama.cpp
Sep 27, 2026
Summary
User reports Qwen3.6-35B-A3B at 706.0 t/s prompt processing on an AMD Radeon RX 9070 XT 16GB under ROCm.
Setup is llama.cpp with Q4_K_M GGUF, flash attention enabled, -ncmoe 40 and -ub 512, on a Ryzen 9 5950X with 64 GB DDR4-3200.
The same configuration under Vulkan reached 333.4 t/s prompt and 29.3 t/s generation; the user retired Vulkan in favour of ROCm.