Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 5080 · Strata · 65,536 ctx
Oct 3, 2026
Use cases
coding
Summary
User reports Qwen3.8-Flash-Next at 110.6 t/s generation and 2773 t/s prefill on two RTX 5080s with a Ryzen 9 9900X, using a custom Strata fork.
Setup is Strata with UD-Q4_K_XL, 64K context, INT8 K/V, 8K prompt chunk, MTP on and expert adaptation on. Upstream v0.1.38 reached 78.1 t/s generation and 2016 t/s prefill with layer split, and 70.3 t/s generation and 1522 t/s prefill with the host peer tier.
With static caches custom was about 26% ahead of upstream layer split (55 vs 44 t/s). Letting the second GPU compute experts during prompt processing doubled custom's 32K prefill (1366 to 2731 t/s), and a 16K chunk brought it to 3724 t/s. Putting the expert tier on the x8 PCIe card improved 4K prefill by 58% and 32K by 34%. The saved answer came from the custom fork, so it may favor its drafts.