Qwen3.8 125B (6B active) Flash-Next
on 2× NVIDIA CMP 170HX 64GB (unlocked) · vLLM · 262,144 ctx
Sep 27, 2026
Use cases
codinglong-contexttool-usemultilingual
Summary
User reports Qwen3.8-Flash-Next at 101.3 t/s decode on Japanese prose and 170.0 t/s on code, single stream, on two NVIDIA CMP 170HX cards.
Setup is vLLM with W4A16 weights, BF16 KV cache, MTP k=4 speculative decoding, 262,144-token context, and expert parallel across the two cards. The PLE n-gram table was converted to FP8 locally to fit in 92 GiB of host RAM.
Aggregate throughput reaches 392 t/s at 4 concurrent requests. Prefill measures 3,191 t/s at 6,954 tokens and 3,261 t/s at 27,853 tokens. A 200,087-token prompt returned the planted code with 70.5 s time to first token and 69.0 t/s decode at depth. MTP is worth about 1.6x on prose and 2.6x on code. The user notes xhigh reasoning effort spent the whole budget without answering in 5 of 6 runs.