Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 4090 · Strata · 32,768 ctx
Oct 5, 2026
Use cases
codingagenticvisiontool-use
Summary
User reports Qwen3.8-Flash-Next at 120.10 tok/s decode at 32K context on one RTX 4090, with 112.55 tok/s at 128K and roughly 4,700-4,800 tok/s on uncached 32K prefill.
Setup is Strata with GSQ IQ3_S GGUF weights, int8 KV cache and a 32K resident window, MTP speculation with four draft tokens, 11 pool workers and a 0.35 PCIe fraction; all experts and the IQ4_NL ngram table stay in RAM, so the run is partially CPU-offloaded. The host is a Ryzen 7900 with 192 GB RAM at a 280 W GPU limit.
The 120.10 tok/s figure is the median of launch medians for the selected 24-expert-cache profile, a +0.88% change over the 96-swap baseline; the 128K figure is +4.26%. Earlier matched decode cells give 110.8 tok/s median for the text profile and 97.65 tok/s for the GPU vision profile, and a matched A/B isolates GPU vision residency at 114.85 versus 104.25 tok/s (-9.2%). The 397B partial-offload probe reached 12.42 tok/s with quality unqualified.