Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 3090 · Strata · 262,144 ctx
Oct 6, 2026
Use cases
agenticvisionlong-context
Summary
User reports Qwen3.8-Flash-Next at 116 t/s decode and 2,908 t/s prompt read on two RTX 3090s with NVLink.
Setup is a modified Strata build (v0.1.38 plus 29 commits) with IQ3_S GGUF at 262K context, greedy, median of 3, on a Ryzen 9 3950X with 121 GB RAM. The second card acts as a peer expert tier read over NVLink; the pair holds about 19,500 of 24,576 IQ3_S experts with a 0.98 hit rate on real agent chats.
Decode gains from the second card are modest (+10%) because the verify window is latency-bound on the primary. Sampled with official settings it sits around 102-104 t/s, a 1.5K prompt reads at ~1,670 t/s, and switching chats takes 1 s. One stream at a time; answers are word-for-word identical to the one-card run.