llamaperf
Oct 6, 2026
Throughput
26.2 t/s gen
Quant
NVFP4 (NVFP4)
KV cache
turbo4
System RAM
254 GB
VRAM reported
32 GB

Summary

User reports Qwen3.8-Flash-Next at 26.2 tok/s single-stream on an RTX 5090 32 GB with 254 GiB system RAM. Setup is FreeToken 0.1.3 with NVFP4 routed experts streamed from system RAM (about 120 GB pinned for offloaded experts) and turbo4 KV cache; MTP is off because verification cost more than direct generation. Aggregate throughput reaches 47.5 tok/s at concurrency 2, 66.9 at 4, 64.6 at 6 and 68.8 at 8, topping out around 65-77 tok/s, limited by PCIe bandwidth and per-step expert fetch volume.