Qwen3.8 27B Swift-1.5
on AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx
Sep 27, 2026
Use cases
tool-uselong-context
Summary
User reports Qwen3.8 27B (Swift-1.5-GSQ-RCO IQ3_S with MTP) at 43.7 tok/s decode and 700.8 t/s prefill on an RX 9070 XT 16GB.
Setup is llama.cpp (server-vulkan b11176) with IQ3_S weights and a q5_0-q4_1 KV cache at 65536 context, one card.
Decode is the mean of 3 runs of a fixed 512-token generation; prefill is one cold 30,461-token prompt. At 131072 context the same Vulkan launch gives 37.4 tok/s decode and 701.6 t/s prefill, while the pinned ROCm image (b10884) cannot load at 128K and decodes 33.5 tok/s at 65536. The user attributes the ROCm failure to a missing FlashAttention vector kernel for the q5_0/q4_1 K/V pair, which promotes both to f16 and exhausts the 16GB card.