Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 3090 · FreeToken
Sep 22, 2026
Summary
User reports Qwen3.8-Flash-Next at 48.0 t/s (median 52.4) on 2x RTX 3090 with tensor parallel 2.
Setup is a fork of FreeToken with the NVFP4 model converted to FTW, 512 experts offloaded, max prefill length 8192, greedy decoding and MTP off.
With MTP on, greedy dropped to 21.5 t/s (median 22.4, accept rate ~56%), so MTP is a net loss on this config; under matched eager conditions it was ~1.5x plain (17.2 vs 11.4).