llamaperf

Qwen3.8 27B

on 2× AMD RX 9070 XT 16GB · llama.cpp · 98,304 ctx

Sep 27, 2026
Throughput
24.7 t/s gen · 2061.1 t/s pp
Quant
Q4_K_M (GGUF)
KV cache
f8_e4m3
System RAM
64 GB
VRAM reported
16 GB

Use cases

codingagentictool-usevisionlong-context

Summary

User reports Qwen3.8-27B at 2061.11 t/s prompt and 24.70 t/s decode on two Radeon RX 9070 XT 16 GB cards under ROCm 7.2 on Windows. Setup is a llama.cpp fork with Q4_K_M GGUF, f8_e4m3 KV cache, FlashAttention, batch 8192 / ubatch 1024, layer split across both GPUs, one server slot, and cold prompt processing at 98,304 context (L3 lane). With MTP n3 the same L1 lane gives 1817.81 t/s prompt and 36.73 t/s decode at 46.8% acceptance. Vulkan rows are lower on prompt processing (L1 1569.17, L2 1470.80, L3 1243.99 t/s) and lead only the L1 decode rows. Linux ROCm 10 results reach 2253.06 t/s prompt with MXFP4-requant and 50.13 t/s decode with Q4_K_M plus MTP n3.