llamaperf

Qwen3.8 27B

on AMD RX 9070 XT 16GB · llama.cpp · 65,536 ctx

Tone: positive
Sep 29, 2026
Throughput
42.0 t/s gen
VRAM reported
16 GB

Use cases

coding

Summary

User reports Qwen 3.8 27B at 42 t/s on an AMD RX 9070 XT 16GB. Setup is llama.cpp with MTP2 speculative decoding and a 64K context, using about 1 GB of working RAM during evaluation. MTP2 delivered +41% to +53% throughput depending on the quantization, but the user notes the fastest configuration was not necessarily best for reasoning quality.