Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 3090 · Strata · 256,000 ctx
Oct 3, 2026
Use cases
codingtool-uselong-context
Summary
User reports Qwen3.8 Flash-Next running on an RTX 3090 24GB with 128GB system RAM, reaching about 1650 t/s prompt processing and 38 to 61 t/s generation depending on context under Strata.
Setup is Strata with Unsloth UD-Q3_K_XL quant, fp16 KV cache, 256k context, speculative decoding with 4 draft tokens, and an expert cache of 6517 slots (~14GB). The same model on llama.cpp master gave up to 700 t/s PP and 23 t/s TG.
Generation is a range: 38 t/s at 182k context and about 61 t/s at short context. The user notes the quant Strata recommends by default was faster but produced minor errors and lower quality, so Unsloth's was chosen for accuracy.