Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 3090 · Strata
Oct 5, 2026
Summary
User reports Qwen3.8 Flash-Next at Q8 running at 34 t/s generation and up to 520 t/s prefill on 2x RTX 3090 with 256GB of system RAM.
Setup uses the Strata engine with MTP set to 5; most of the model sits in system RAM rather than GPU memory.
For comparison, the user says llama.cpp without MTP gives 13 t/s generation and 115 t/s prefill on the same setup, and notes MTP usually does not help when most of the model is in system RAM.