Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 5070 Ti Laptop 12GB · Strata · 131,072 ctx
Sep 30, 2026
Summary
User reports Qwen3.8 Flash-Next at 51 t/s generation and 1500 t/s prompt processing on an RTX 5070 Ti Laptop 12GB with 64GB RAM.
Setup is the Strata engine with an IQ3_XXS GGUF quant and 8-bit KV cache, using 131k context. The model was partially offloaded, using 11GB VRAM and 56GB system RAM.
The generation speed was measured at 43k context depth. The user compares this to stock llama.cpp, which reached 23 t/s generation and 100 t/s prompt processing with the same quant.