Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 3090 · Strata
Oct 4, 2026
Summary
User reports Qwen3.8 Flash-Next at 80-110 tok/s on 2x RTX 3090 with 96GB DDR5 system RAM.
Setup uses the Strata engine with a Q4_K_XL GGUF quant from bitlamas, at 2500 prompt tokens.
The user calls it the biggest thing since Qwen 3.8 27B, offering the speed of Qwen 3.5 35B-A3B with higher intelligence than 3.8 27B.