Qwen3.8 27B
on 2× AMD RX 9070 XT 16GB · llama.cpp · 98,304 ctx
Sep 27, 2026
Use cases
codingagentictool-usevisionlong-context
Summary
User reports Qwen3.8-27B at 2061.11 t/s prompt and 24.70 t/s decode on two Radeon RX 9070 XT 16 GB cards under ROCm 7.2 on Windows.
Setup is a llama.cpp fork with Q4_K_M GGUF, f8_e4m3 KV cache, FlashAttention, batch 8192 / ubatch 1024, layer split across both GPUs, one server slot, and cold prompt processing at 98,304 context (L3 lane).
With MTP n3 the same L1 lane gives 1817.81 t/s prompt and 36.73 t/s decode at 46.8% acceptance. Vulkan rows are lower on prompt processing (L1 1569.17, L2 1470.80, L3 1243.99 t/s) and lead only the L1 decode rows. Linux ROCm 10 results reach 2253.06 t/s prompt with MXFP4-requant and 50.13 t/s decode with Q4_K_M plus MTP n3.