Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 5090 · FreeToken · 229,376 ctx
Sep 20, 2026
Summary
User reports Qwen3.8 Flash-Next at 50 t/s generation and 2300 t/s prompt processing on a single RTX 5090 with 128 GB DDR5.
Setup is FreeToken with the nvidia/Qwen3.8-Flash-Next-NVFP4 model converted to FreeToken format, expert caching enabled, max sequence length 229376, and max running requests 2.
Generation speed ranges from 40 to 60 t/s depending on cache effectiveness. User notes llama.cpp lacks MoE expert caching, and FreeToken is early with no KV quantization or MTP yet.