Qwen3.8 125B (6B active) Flash-Next
on 4× NVIDIA CMP 170HX 40GB (unlocked) · llama.cpp · 262,144 ctx
Sep 27, 2026
Use cases
long-contextagentic
Summary
User reports Qwen3.8-Flash-Next at 96.9 tok/s single-request decode on four NVIDIA CMP 170HX 40GB cards.
Setup is llama.cpp with UD-Q4_K_XL weights, q8_0 KV cache, MTP speculative decoding (draft length 4, GPU sampling), 262K context, and the SM clock pinned at 1410 MHz.
The same configuration reaches 87.7 tok/s at 70K context; three concurrent 3K requests give 34-37 tok/s each (about 100 tok/s aggregate). Earlier revisions of the fork measured 64-73 tok/s at 2K and 58-70 tok/s at 70K, against a baseline fork at 46-55 tok/s and 27-45 tok/s respectively.