llamaperf

Qwen3.8 27B Huihui-Abliterated

on NVIDIA RTX 5090 · SGLang · 253,952 ctx

Tone: mixed
Sep 18, 2026
Throughput
100.0 t/s gen
Quant
NVFP4 (NVFP4)
KV cache
fp8_e4m3
VRAM reported
32 GB

Use cases

multilinguallong-contextvision

Summary

User reports Qwen3.8 27B at about 100 t/s on a single RTX 5090, but only about 82k tokens of usable context despite configuring 253,952. Setup is SGLang serving the Huihui-Abliterated NVFP4 MTP VL build with an fp8_e4m3 KV cache, NEXTN speculative decoding at 3 steps and 4 draft tokens, flashinfer attention, chunked prefill of 1024, and a 16 GB hierarchical cache. User asks how to tune the setup for more context without losing multimodal capability.