Qwen3.6 35B (3B active)
AMD RX 9060 XT 16GB · llama.cpp · 40,960 ctx
- reported speed:
- 66.0 tokens/s generation
- quant:
- UD-Q4_K_XL (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.6-35B-A3B at 66.04 tokens/s on an AMD Radeon RX 9060 XT 16 GB with 32 GB of system RAM. Setup is llama.cpp (Vulkan backend) with UD-Q4_K_XL GGUF weights, q8_0 KV cache, 40k context, and MTP speculative decoding drafting up to 3 tokens per step. The 35B MoE model does not fit in 16 GB of VRAM, so the expert weights of the first 20 layers run on the CPU. Run on a Ryzen 7 5800X3D desktop under SteamOS, single slot, with flash attention and prefix cache reuse enabled.