Qwen3.8 125B (6B active) Flash-Next
T4 16GB · ik_llama.cpp · 262,144 ctx
- reported speed:
- 17.6 tokens/s generation · 159.6 tokens/s prompt processing
- quant:
- UD-Q4_K_XL (gguf)
- kv:
- Q8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Qwen3.8-Flash-Next, 180B total with 6B active across 512 experts, generating 17.6 t/s on short prompts and 16.1 t/s at 12.5K context. Non-expert weights run on a T4 using 4606 MiB, with the experts held in host RAM. Prompt processing reaches 159.6 t/s on a cold 12.5K prompt. User reports solid refactor and coding performance, and finds it less verbose than Opus.