Qwen3.8 27B
on NVIDIA RTX 5090 · SGLang · 147,456 ctx
Sep 30, 2026
Summary
User reports Qwen3.8-27B NVFP4 at 152 tok/s decode on one RTX 5090 under WSL2, with MTP speculative decoding and host embedding enabled.
Setup is SGLang v0.5.20 with NVFP4 weights, 147,456-token context, on a Ryzen 9 9950X3D with 96 GB RAM.
Plain 27B without speculation reached 93 tok/s at one stream (89.7-96.0) and 456 tok/s total at 5 streams; with MTP, 5 streams totalled 753 tok/s. Prefill of a 6K-token prompt was 13.4-15.0K tok/s plain and 9.2-12.4K tok/s with MTP. Energy was 2.91 J/token plain and 1.72 J/token with MTP.