Qwen3.5 9B
on AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx
Sep 27, 2026
Summary
User benchmarks Qwen3.5-9B at 62 t/s on an AMD Radeon RX 9070 XT (RDNA4/gfx1201) using llama-server with the Vulkan backend.
Setup is llama.cpp with GGUF Q6_K_XL weights at 65536 tokens of context, single request, no speculation. The same model under vLLM ROCm 7.2 with FP8 weights reached 48 t/s, which the user attributes to vLLM lacking native gfx1201 kernel support and falling back to FP32 dequantization.
The user reports llama-server Vulkan is 29% faster on this hardware and recommends it over vLLM until RDNA4 support lands upstream.