Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 3090 · Strata
Oct 7, 2026
Use cases
codingagentic
Summary
User reports Qwen3.8-Flash-Next at 97 t/s decode on 2x RTX 3090 with Strata, with about half the experts in system RAM.
Setup is Strata v0.1.40.1 with unsloth UD-Q4_K_XL GGUF, a layer split, speculative decoding, and 121 GB DDR4 on a Ryzen 9 3950X.
The 97 t/s is at a 95% expert hit rate; shrinking the cache to 90% and 81% hit rates gave 84 t/s and 68 t/s. The user estimates +34% decode from eliminating misses on their box and about +75% headroom at a single-24GB-card share, and reports +16% decode from --pcie-frac 0.2 --spec-min-p 0.7 over about 280 runs. The prefetch feature is not implemented yet.