llamaperf

Qwen3.6 35B (3B active)

on AMD RX 9060 XT 16GB · llama.cpp · 40,960 ctx

Oct 6, 2026
Throughput
66.0 t/s gen
Quant
UD-Q4_K_XL (GGUF)
KV cache
q8_0
System RAM
32 GB
VRAM reported
16 GB

Summary

User reports Qwen3.6-35B-A3B at 66.04 tokens/s on an AMD Radeon RX 9060 XT 16 GB with 32 GB of system RAM. Setup is llama.cpp (Vulkan backend) with UD-Q4_K_XL GGUF weights, q8_0 KV cache, 40k context, and MTP speculative decoding drafting up to 3 tokens per step. The 35B MoE model does not fit in 16 GB of VRAM, so the expert weights of the first 20 layers run on the CPU. Run on a Ryzen 7 5800X3D desktop under SteamOS, single slot, with flash attention and prefix cache reuse enabled.