Qwen3.8 125B (6B active) Flash-Next
on RTX 2060 Laptop 6GB · Strata · 50,000 ctx
Oct 3, 2026
Summary
User reports Qwen 3.8 Flash Next at 10 t/s decode and almost 100 t/s prefill on an RTX 2060 Laptop 6GB with 32 GB RAM.
Setup is Strata with q2_0 quant and Q4_0 KV cache at 50k context, using draft-vocab en and 100 MiB VRAM reserve.
User notes prefill is the bottleneck for large prompts, and that Qwen A35B A3B has around 5x faster prefill and allows 100K context without KV quantization.