llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 5090 · FreeToken · 229,376 ctx

Tone: mixed
Sep 20, 2026
Throughput
50.0 t/s gen · 2300.0 t/s pp
Quant
NVFP4
System RAM
128 GB

Summary

User reports Qwen3.8 Flash-Next at 50 t/s generation and 2300 t/s prompt processing on a single RTX 5090 with 128 GB DDR5. Setup is FreeToken with the nvidia/Qwen3.8-Flash-Next-NVFP4 model converted to FreeToken format, expert caching enabled, max sequence length 229376, and max running requests 2. Generation speed ranges from 40 to 60 t/s depending on cache effectiveness. User notes llama.cpp lacks MoE expert caching, and FreeToken is early with no KV quantization or MTP yet.