llamaperf

Qwen3.6 35B (3B active)

on AMD RX 9070 XT 16GB · llama.cpp · 32,768 ctx

Tone: mixed
Oct 5, 2026
Throughput
62.0 t/s gen
System RAM
64 GB
VRAM reported
16 GB

Use cases

codingagentictool-usesummarization

Summary

User reports Qwen3.6 35B A3B at 62 t/s on an AMD RX 9070 XT 16GB, with vLLM on ROCm reaching 48 t/s in the same comparison. Setup is llama.cpp with the Vulkan backend, 32k context, on a Ryzen 7 9800X3D with 64GB DDR5; the MoE model spills part of its weights to system RAM. User notes the setup is early and numbers are not settled, that 16GB VRAM rules out big dense models, and that the local model needs supervision.