Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX Pro 6000 Blackwell · NInfer · 512 ctx
Oct 6, 2026
Summary
User reports Qwen 3.8 Flash Next at 172.0 tok/s decode without speculative decoding on an RTX6000 96GB, at 512 context with non-expert weights downsampled from 16-bit to 8-bit.
Setup is NInfer6000, the user's fork of NInfer, with NVFP4 quants and prefill at 5,905 tok/s on 512 tokens and 13,908 tok/s on 8,192 tokens.
With MTP3 and --lm-head-draft the same setup reaches 274.8 tok/s at 512 and 401.3 tok/s at 8K. Keeping non-experts at 16-bit gives 118.0 tok/s without speculative decoding.