Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 4090 · Strata · 131,072 ctx
Oct 5, 2026
Summary
User reports Qwen3.8-Flash-Next at 106.1 tok/s whole-request on an RTX 4090 24GB.
Setup is Strata with IQ2_XS, 128K context, 8-bit KV cache, MTP on, KV streaming on, images on, speed projection off.
The IQ3_XXS tier ran at 98.1 tok/s whole-request with ~120 tok/s peak decode and 131 tok/s prefill; the user notes it was 3.7x longer and qualitatively deeper for only 8% less speed. Both tiers match the README's 24 GB card estimate of 100-140 t/s.