Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 5090 · Strata · 262,144 ctx
Oct 3, 2026
Summary
User reports Qwen3.8-Flash-Next at 1,796 tok/s prompt and 128-134 tok/s decode on an RTX 5090 32 GB plus an RTX 4070 Ti SUPER 16 GB.
Setup is Strata with IQ3_XXS GGUF, int8 KV cache, 32K cells resident per layer, MTP spec 4, vision on, and 262,144 context; the 76 GB model runs with 31 GB of system RAM and experts streamed from an mmap'd GGUF.
Decode is a range of 128-134 tok/s median on real sampling and 95-108 greedy; stock 0.1.33 with a layer split at 36 gave 890 tok/s on an 80K prompt and about 110 tok/s decode. Three patches cut major faults from 24 million to 14 thousand and chat switching from 20-48 s to 0.6-1.2 s.