Qwen3.6 35B (3B active)
AMD RX 9070 XT 16GB · llama.cpp · 32,768 ctx
- reported speed:
- 62.0 tokens/s generation
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.6 35B A3B at 62 t/s on an AMD RX 9070 XT 16GB, with vLLM on ROCm reaching 48 t/s in the same comparison. Setup is llama.cpp with the Vulkan backend, 32k context, on a Ryzen 7 9800X3D with 64GB DDR5; the MoE model spills part of its weights to system RAM. User notes the setup is early and numbers are not settled, that 16GB VRAM rules out big dense models, and that the local model needs supervision.