llamaperf
Sep 29, 2026
Throughput
50.5 t/s gen
Quant
NVFP4 (NVFP4)
System RAM
96 GB

Summary

User reports Qwen3.8-Flash-Next Uncensored NVFP4 at a median 50.5 t/s generation on one RTX 5090, measured at the maximum input of 259,601 tokens. Setup is a modified FreeToken engine following the llama-split-bench protocol, with 96 GB system RAM and 1,000 tokens generated per test. Across 3 rounds and 36 completed measurements, the median over 9 short practical prompts was 55.4 t/s.