Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 5070 · Strata · 131,072 ctx
Sep 24, 2026
Summary
User reports Qwen3.8-Flash-Next at 44.8 t/s generation and 414 t/s prompt processing on a 12GB RTX 5070 at 128K context.
Setup is a custom Strata inference engine with the IQ3_XXS GGUF quant, 64GB DDR5-5600 and a Ryzen 5 7600 on Windows. The engine is CUDA-only and requires 47GB minimum in RAM+VRAM for this quant.
A Q2_0 quant reached 65.1 t/s generation and 543 t/s prompt processing, and IQ2_XS reached 52.0 t/s generation and 472 t/s prompt processing. The user previously measured 15 t/s output and 100-120 t/s prompt processing with llama.cpp on the same hardware.