Oct 7, 2026
Summary
User reports Qwen3.8 Flash-Next at 55 t/s on an RTX 4080 with 128 GB DDR5 RAM, using Strata with Q3 weights and Q8 KV cache at roughly 120-150k context.
Setup is Strata with the 3060 left unused; adding the second GPU drops generation to 45 t/s and prompt processing to 250 t/s from 590 t/s.
With Unsloth Studio the same model reached 27 t/s on the 4080 alone, 23 t/s with both GPUs, and 16 t/s with both GPUs on an older PCIe 4.0 x1 link. The user asks whether the 3060 is counter-productive alongside the 4080.