llamaperf

Qwen3.8 125B (6B active) Flash-Next

on 4× RTX 5060 Ti · Strata · 262,144 ctx

Tone: positive
Oct 7, 2026
Throughput
104.0 t/s gen
Quant
IQ3_XXS (GGUF)

Summary

User reports Qwen3.8 Flash-Next at 104 tok/s on a mixed 4-GPU rig (RTX 5060 Ti, RTX 3090, 2x RTX 3060). Setup is Strata with IQ3_XXS at 262k context; the 125B MoE model spills experts over PCIe, leaving about 10.7 GiB of the 63.9 GiB VRAM free. The user explains the unused VRAM is headroom: 0.1.40.1 caches 24% fewer resident experts than 0.1.39 (17,515 vs 23,168) while cache hit rate only fell from 99.7% to 98.1%, and decode rose from about 87 tok/s to 104 tok/s. The 104 tok/s run used a 250 W power cap versus 370 W for the earlier runs.