Qwen3.8 27B
on AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx
Sep 29, 2026
Use cases
coding
Summary
User reports Qwen 3.8 27B at 42 t/s on an AMD RX 9070 XT 16GB.
Setup is llama.cpp with MTP2 speculative decoding and a 64K context, using about 1 GB of working RAM during evaluation.
MTP2 delivered +41% to +53% throughput depending on the quantization, but the user notes the fastest configuration was not necessarily best for reasoning quality.