Qwen3.8 125B (6B active) Swift-1.5
on 2× NVIDIA RTX 5080 · Strata · 262,144 ctx
Oct 3, 2026
Use cases
coding
Summary
User reports Qwen3.8-Flash-Next (Swift 1.5 IQ2_XS) at 143 tok/s code generation on an RTX 5080 + RTX 4060 Ti with 32 GB system RAM.
Setup is a dual-GPU fork of the Strata engine at 256K context, with experts that don't fit in VRAM read from the SSD and a fine-tuned draft layer for speculative decoding.
The same machine with the 5080 alone reached 29 tok/s code and 333 tok/s on a 32K prompt; the fork with both cards reached 1,940 tok/s on the 32K prompt. In real use with the Pi coding agent it peaks above 200 tok/s (209 so far) and a 150K-token context reads at about 1,850 tok/s. Perplexity was 7.24 vs upstream's 7.28.