llamaperf

Qwen3.8 27B Huihui-Abliterated

on NVIDIA RTX 5090 · NInfer · 196,608 ctx

Tone: positive
Sep 17, 2026
Throughput
205.0 t/s gen · 9000.0 t/s pp
Quant
NVFP4 (NVFP4)
KV cache
fp8
VRAM reported
32 GB

Use cases

codingcreative-writinglong-contextvision

Summary

User reports Qwen3.8 27B Huihui abliterated NVFP4 at 205 t/s decode on a single RTX 5090 32GB, with prefill around 9,000 t/s at 8K context. Setup is NInfer on Windows 11 + WSL2 with fp8 KV cache at 196,608 context, MTP speculative decoding with 3 draft tokens and --lm-head-draft, vision enabled, weights about 19.7 GB. Decode varies with MTP acceptance: 253 t/s on predictable text, 205 t/s on coding, about 120 t/s on creative prose, and 60-80 t/s with MTP off. At 128K context decode is about 115 t/s and prefill about 4,500 t/s.