Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 3090 · Strata · 262,144 ctx
Oct 7, 2026
Summary
User reports Qwen3.8-Flash-Next at 43.9 tok/s generation and 1817 tok/s prompt processing on an RTX 3090 + RTX A4000 at 32k context.
Setup is Strata with UD-Q6_K_XL GGUF and int8 KV cache, 262144 native context, speculative decoding with MTP, experts offloaded to CPU.
At 200k context Strata Q6 does 48.8 tok/s versus llama.cpp's 10.6 tok/s, and TTFT drops from 595 s to 118 s. Q6 is 12-31% faster than Q8 and 19 GB smaller.