Qwen3.8 125B (6B active) Flash-Next
4× NVIDIA CMP 170HX 40GB (unlocked) · llama.cpp · 262,144 ctx
- reported speed:
- 96.9 tokens/s generation
- quant:
- UD-Q4_K_XL (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-Flash-Next at 96.9 tok/s single-request decode on four NVIDIA CMP 170HX 40GB cards. Setup is llama.cpp with UD-Q4_K_XL weights, q8_0 KV cache, MTP speculative decoding (draft length 4, GPU sampling), 262K context, and the SM clock pinned at 1410 MHz. The same configuration reaches 87.7 tok/s at 70K context; three concurrent 3K requests give 34-37 tok/s each (about 100 tok/s aggregate). Earlier revisions of the fork measured 64-73 tok/s at 2K and 58-70 tok/s at 70K, against a baseline fork at 46-55 tok/s and 27-45 tok/s respectively.