llamaperf

Qwen3.8 27B

on AMD RX 7600 XT 16GB · llama.cpp · 163,840 ctx

Tone: positive
Sep 23, 2026
Throughput
18.0 t/s gen · 141.0 t/s pp
Quant
IQ3_XXS (GGUF)
KV cache
q8_0
MTP (Multi-Token Prediction)
off
System RAM
32 GB
VRAM reported
16 GB

Use cases

coding

Summary

User reports Qwen3.8 27B at 18 t/s decode and 141 t/s prefill on a 1.8k-token prompt, running on an AMD Radeon RX 7600 XT 16 GB at 163,840 context. Setup is llama.cpp llama-server build 10480 with the unsloth Qwen3.8-27B-UD-IQ3_XXS GGUF (10.2 GiB), q8_0 KV cache, flash attention on, and Vulkan (RADV) backend. The model is a hybrid architecture where only 16 of 64 layers keep a full KV cache, so KV is about 5.3 GiB at this context. With MTP on the same model runs about 24 t/s at 98k context and 39 t/s at 124k. The user notes 160k is a VRAM-math ceiling, not a quality claim, and that decode drops to about 6 t/s if the GPU spills to system RAM.