Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 3090 · Strata · 131,072 ctx
Oct 5, 2026
Summary
User reports running DeepSeek V4 Flash with Strata on an RTX 3090 alone at 1700 t/s prefill and 37-40 t/s generation, and on both an RTX 3090 and RTX 5070 Ti at 1850 t/s prefill and 45-50 t/s generation, at 131K context.
Setup is Strata with UD-Q4_K_XL quant, 96GB DDR4 system RAM, RTX 5070 Ti on PCIe x16 and RTX 3090 on PCIe x4.
User compares against Unsloth Studio on the same model at 131K context, which gave 100 t/s prefill and 14-16 t/s generation, with a 100K conversation taking over 10 minutes to process. User notes Strata initially lacked prefix caching but now supports it, and that the same generation speed matches Qwen3.8-27B-UD-Q8_K_L on both GPUs.