Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 98,304 ctx
Oct 5, 2026
Use cases
coding
Summary
User reports Qwen3.8-Flash-Next at 30.5 t/s generation and 250.4 t/s prompt processing on 2x RTX 5060 Ti with 32GB system RAM.
Setup is llama.cpp with an IQ1_M GGUF (27.58 GB), 98304 context, q8_0 KV cache, tensor split 1,1, and 8 MoE layers offloaded to CPU.
User asks whether their settings are correct and what config others run for this model on this hardware.