llamaperf

Qwen3.5 9B

on AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx

Tone: positive
Sep 27, 2026
Throughput
62.0 t/s gen
Quant
Q6_K_XL (GGUF)

Summary

User benchmarks Qwen3.5-9B at 62 t/s on an AMD Radeon RX 9070 XT (RDNA4/gfx1201) using llama-server with the Vulkan backend. Setup is llama.cpp with GGUF Q6_K_XL weights at 65536 tokens of context, single request, no speculation. The same model under vLLM ROCm 7.2 with FP8 weights reached 48 t/s, which the user attributes to vLLM lacking native gfx1201 kernel support and falling back to FP32 dequantization. The user reports llama-server Vulkan is 29% faster on this hardware and recommends it over vLLM until RDNA4 support lands upstream.