Qwen3.8 27B Huihui-Abliterated
on NVIDIA RTX 5090 · NInfer · 196,608 ctx
Sep 17, 2026
Use cases
codingcreative-writinglong-contextvision
Summary
User reports Qwen3.8 27B Huihui abliterated NVFP4 at 205 t/s decode on a single RTX 5090 32GB, with prefill around 9,000 t/s at 8K context.
Setup is NInfer on Windows 11 + WSL2 with fp8 KV cache at 196,608 context, MTP speculative decoding with 3 draft tokens and --lm-head-draft, vision enabled, weights about 19.7 GB.
Decode varies with MTP acceptance: 253 t/s on predictable text, 205 t/s on coding, about 120 t/s on creative prose, and 60-80 t/s with MTP off. At 128K context decode is about 115 t/s and prefill about 4,500 t/s.