llamaperf

Qwen3.8 27B Swift

on 2× AMD RX 9070 XT 16GB · llama.cpp · 131,072 ctx

Tone: mixed
Sep 21, 2026
Throughput
50.0 t/s gen
Quant
Q6 (GGUF)
KV cache
Q8
System RAM
64 GB

Summary

User reports Qwen 3.8 Swift at a ceiling of 50 t/s on an RX 9070 XT 16GB and a Radeon AI PRO R9700 32GB. Setup is llama.cpp with the Vulkan backend, a Q6 quant, 131K context, and Q8 KV cache for both K and V, with a DFlash2 drafter. User asks whether others get more than 50 t/s on AMD 9070 XT or R9700 hardware with a usable context size, and whether a different OS could add 30+ t/s.