llamaperf
Sep 23, 2026
Throughput
116.0 t/s gen · 6800-7950 t/s pp
Quant
W4A16
System RAM
128 GB
VRAM reported
64 GB

Summary

User reports Qwen3-Next-80B at 116 t/s single-stream decode on a single NVIDIA CMP 170HX with 64 GB HBM2e, at a 150 W cap. Setup is vLLM 0.27.1 with W4A16 weights (40.9 GB), torch 2.13.0+cu130, CUDA 13.0, Ubuntu 26.04 LTS, on an AMD Ryzen Threadripper PRO 3945WX with 128 GB DDR4 ECC. Prefill over about 8.9k tokens measured 6800 to 7950 t/s; power draw 137 to 145 W. Aggregate throughput at 8 concurrent requests was 352 t/s, saturating at 4 slots. Also measured on the same card for context: Ornith-1.5-35B FP8 at 122.5 t/s and Qwen3.8-27B W4A16 with DFlash2 at 127 t/s single-stream.