Qwen3.8 125B (6B active) Flash-Next Uncensored
on NVIDIA RTX 5090 · FreeToken · 259,601 ctx
Sep 29, 2026
Summary
User reports Qwen3.8-Flash-Next Uncensored NVFP4 at a median 50.5 t/s generation on one RTX 5090, measured at the maximum input of 259,601 tokens.
Setup is a modified FreeToken engine following the llama-split-bench protocol, with 96 GB system RAM and 1,000 tokens generated per test.
Across 3 rounds and 36 completed measurements, the median over 9 short practical prompts was 55.4 t/s.