llamaperf

Qwen3.8 27B

on AMD RX 9070 XT 16GB · llama.cpp · 131,072 ctx

Sep 28, 2026
Throughput
50.0 t/s gen
Quant
Q6 (GGUF)
KV cache
Q8
System RAM
64 GB

Summary

User reports 50 t/s for Qwen 3.8 27B on an RX 9070 XT 16GB. Setup is llama.cpp (Vulkan build) with a Q6 quant, 131072 context, Q8 K cache and Q4 V cache, and DFlash2 speculative decoding with 7 max drafts. The rig also has an R9700 AI Pro 32GB and 64GB DDR5 6000; the user asks what speeds and flags others get, especially on Windows/AMD.