Qwen3.6 35B (3B active)
on AMD RX 9070 XT 16GB · llama.cpp · 32,768 ctx
Oct 5, 2026
Use cases
codingagentictool-usesummarization
Summary
User reports Qwen3.6 35B A3B at 62 t/s on an AMD RX 9070 XT 16GB, with vLLM on ROCm reaching 48 t/s in the same comparison.
Setup is llama.cpp with the Vulkan backend, 32k context, on a Ryzen 7 9800X3D with 64GB DDR5; the MoE model spills part of its weights to system RAM.
User notes the setup is early and numbers are not settled, that 16GB VRAM rules out big dense models, and that the local model needs supervision.