Qwen3.8 125B (6B active) Flash-Next
on NVIDIA V100 16GB · Strata · 1,036 ctx
Oct 5, 2026
Summary
User reports Qwen3.8-Flash-Next at 52.0 t/s decode and 398.6 t/s prefill on a single Tesla V100-PCIE-16GB (PCIe Gen3 x16).
Setup is the Strata-V100 fork with Q2_0 weights, int8 KV cache, 262,144-token context, MTP with --spec 8 and draft floor 0.70, a 20/28 layer split, and 700 MiB vision reserve; the machine has a Ryzen 5 3600 and 48 GB DDR4-3200.
Each value is the median of three fresh uncached requests with exactly 256 output tokens. Decode falls from 52.0 t/s at 1K to 36.8 t/s at 256K; the 128K and 256K prompts hit 84 °C and thermally throttled during prefill. A separate before/after run of the integrated build measured prefill gains of +2.44% at 2K, +3.93% at 8K and +0.94% at 32K.