llamaperf

Qwen3.8 27B Huihui-Abliterated

on AMD RX 6900 XT 16GB · llama.cpp · 100,000 ctx

Tone: positive
Sep 26, 2026
Throughput
30-45 t/s gen
Quant
IQ3_XXS (GGUF)
KV cache
q8_0/q4_0
MTP (Multi-Token Prediction)
on
VRAM reported
16 GB

Summary

User reports running Qwen3.8-27b (Huihui abliterated, IQ3_XXS) on an RX 6900 XT 16GB with llama.cpp and ROCm 10. Prefill drops from 400 t/s to 260 t/s at 40k context; decode varies between 30 t/s and 45 t/s. Setup uses 100k context, KV cache q8_0/q4_0, flash attention, ngram-mod, MTP with n=2, and mmproj in system RAM; VRAM fills to 15.7/16 GB.