Qwen3.8 125B (6B active) Flash-Next
on NVIDIA RTX 5090 · FreeToken
Oct 6, 2026
Summary
User reports Qwen3.8-Flash-Next at 26.2 tok/s single-stream on an RTX 5090 32 GB with 254 GiB system RAM.
Setup is FreeToken 0.1.3 with NVFP4 routed experts streamed from system RAM (about 120 GB pinned for offloaded experts) and turbo4 KV cache; MTP is off because verification cost more than direct generation.
Aggregate throughput reaches 47.5 tok/s at concurrency 2, 66.9 at 4, 64.6 at 6 and 68.8 at 8, topping out around 65-77 tok/s, limited by PCIe bandwidth and per-step expert fetch volume.