llamaperf

Qwen3.8 27B Swift-1.5

on AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx

Tone: positive
Sep 27, 2026
Throughput
43.7 t/s gen · 700.8 t/s pp
Quant
IQ3_S (GGUF)
KV cache
q5_0-q4_1
VRAM reported
16 GB

Use cases

tool-uselong-context

Summary

User reports Qwen3.8 27B (Swift-1.5-GSQ-RCO IQ3_S with MTP) at 43.7 tok/s decode and 700.8 t/s prefill on an RX 9070 XT 16GB. Setup is llama.cpp (server-vulkan b11176) with IQ3_S weights and a q5_0-q4_1 KV cache at 65536 context, one card. Decode is the mean of 3 runs of a fixed 512-token generation; prefill is one cold 30,461-token prompt. At 131072 context the same Vulkan launch gives 37.4 tok/s decode and 701.6 t/s prefill, while the pinned ROCm image (b10884) cannot load at 128K and decodes 33.5 tok/s at 65536. The user attributes the ROCm failure to a missing FlashAttention vector kernel for the q5_0/q4_1 K/V pair, which promotes both to f16 and exhausts the 16GB card.