Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 3060 12GB · ik_llama.cpp · 262,144 ctx
Sep 23, 2026
Use cases
codingtool-useagenticlong-context
Summary
User reports Qwen3.8 Flash Next Q4_K_XL at 66.14 t/s prompt processing and 22.09 t/s generation on an RTX 3060 12GB, Intel i7-14700 and 96GB DDR5.
Setup is ik_llama / llama-server on Windows 11 at 262,144 tokens of context with MTP3 and NCMOE48. A fresh OpenCode agent session produced a 10,890-token initial prompt, and MTP draft acceptance for that request was 89.5%.
After the initial ~11K-token prefill, a second short prompt completed in 1.6s with short-generation throughput around 22–26 t/s. At approximately 71K active tokens, generation was around 10 t/s, with an OpenCode session average around 12 t/s.