Qwen3.8 27B Swift
on 2× AMD RX 9070 XT 16GB · llama.cpp · 131,072 ctx
Sep 21, 2026
Summary
User reports Qwen 3.8 Swift at a ceiling of 50 t/s on an RX 9070 XT 16GB and a Radeon AI PRO R9700 32GB.
Setup is llama.cpp with the Vulkan backend, a Q6 quant, 131K context, and Q8 KV cache for both K and V, with a DFlash2 drafter.
User asks whether others get more than 50 t/s on AMD 9070 XT or R9700 hardware with a usable context size, and whether a different OS could add 30+ t/s.