llamaperf
Oct 6, 2026
Throughput
172.0 t/s gen · 5905.0 t/s pp
Quant
NVFP4 (NVFP4)
VRAM reported
96 GB

Summary

User reports Qwen 3.8 Flash Next at 172.0 tok/s decode without speculative decoding on an RTX6000 96GB, at 512 context with non-expert weights downsampled from 16-bit to 8-bit. Setup is NInfer6000, the user's fork of NInfer, with NVFP4 quants and prefill at 5,905 tok/s on 512 tokens and 13,908 tok/s on 8,192 tokens. With MTP3 and --lm-head-draft the same setup reaches 274.8 tok/s at 512 and 401.3 tok/s at 8K. Keeping non-experts at 16-bit gives 118.0 tok/s without speculative decoding.