Qwen3.8 27B Huihui-Abliterated
on AMD RX 6900 XT 16GB · llama.cpp · 100,000 ctx
Sep 26, 2026
Summary
User reports running Qwen3.8-27b (Huihui abliterated, IQ3_XXS) on an RX 6900 XT 16GB with llama.cpp and ROCm 10. Prefill drops from 400 t/s to 260 t/s at 40k context; decode varies between 30 t/s and 45 t/s. Setup uses 100k context, KV cache q8_0/q4_0, flash attention, ngram-mod, MTP with n=2, and mmproj in system RAM; VRAM fills to 15.7/16 GB.