Qwen3.8 125B (6B active) Flash-Next
on 4× RTX 5060 Ti · Strata · 262,144 ctx
Oct 7, 2026
Summary
User reports Qwen3.8 Flash-Next at 104 tok/s on a mixed 4-GPU rig (RTX 5060 Ti, RTX 3090, 2x RTX 3060).
Setup is Strata with IQ3_XXS at 262k context; the 125B MoE model spills experts over PCIe, leaving about 10.7 GiB of the 63.9 GiB VRAM free.
The user explains the unused VRAM is headroom: 0.1.40.1 caches 24% fewer resident experts than 0.1.39 (17,515 vs 23,168) while cache hit rate only fell from 99.7% to 98.1%, and decode rose from about 87 tok/s to 104 tok/s. The 104 tok/s run used a 250 W power cap versus 370 W for the earlier runs.