llamaperf

Qwen3.8 27B

on NVIDIA RTX 5090 · SGLang · 147,456 ctx

Sep 30, 2026
Throughput
152.0 t/s gen
Quant
NVFP4 (NVFP4)
System RAM
96 GB

Summary

User reports Qwen3.8-27B NVFP4 at 152 tok/s decode on one RTX 5090 under WSL2, with MTP speculative decoding and host embedding enabled. Setup is SGLang v0.5.20 with NVFP4 weights, 147,456-token context, on a Ryzen 9 9950X3D with 96 GB RAM. Plain 27B without speculation reached 93 tok/s at one stream (89.7-96.0) and 456 tok/s total at 5 streams; with MTP, 5 streams totalled 753 tok/s. Prefill of a 6K-token prompt was 13.4-15.0K tok/s plain and 9.2-12.4K tok/s with MTP. Energy was 2.91 J/token plain and 1.72 J/token with MTP.