Qwen3.6 35B (3B active)
on AMD RX 9060 XT 16GB · llama.cpp · 40,960 ctx
Oct 6, 2026
Summary
User reports Qwen3.6-35B-A3B at 66.04 tokens/s on an AMD Radeon RX 9060 XT 16 GB with 32 GB of system RAM.
Setup is llama.cpp (Vulkan backend) with UD-Q4_K_XL GGUF weights, q8_0 KV cache, 40k context, and MTP speculative decoding drafting up to 3 tokens per step. The 35B MoE model does not fit in 16 GB of VRAM, so the expert weights of the first 20 layers run on the CPU.
Run on a Ryzen 7 5800X3D desktop under SteamOS, single slot, with flash attention and prefix cache reuse enabled.