Sep 23, 2026
Summary
User reports Qwen3-Next-80B at 116 t/s single-stream decode on a single NVIDIA CMP 170HX with 64 GB HBM2e, at a 150 W cap.
Setup is vLLM 0.27.1 with W4A16 weights (40.9 GB), torch 2.13.0+cu130, CUDA 13.0, Ubuntu 26.04 LTS, on an AMD Ryzen Threadripper PRO 3945WX with 128 GB DDR4 ECC. Prefill over about 8.9k tokens measured 6800 to 7950 t/s; power draw 137 to 145 W.
Aggregate throughput at 8 concurrent requests was 352 t/s, saturating at 4 slots. Also measured on the same card for context: Ornith-1.5-35B FP8 at 122.5 t/s and Qwen3.8-27B W4A16 with DFlash2 at 127 t/s single-stream.