llamaperf

Qwen3.8 125B (6B active) Flash-Next

on NVIDIA RTX 4090 · Strata · 131,072 ctx

Tone: positive
Oct 5, 2026
Throughput
106.1 t/s gen · 147.0 t/s pp
Quant
IQ2_XS (GGUF)
KV cache
8-bit
System RAM
64 GB
VRAM reported
24 GB

Summary

User reports Qwen3.8-Flash-Next at 106.1 tok/s whole-request on an RTX 4090 24GB. Setup is Strata with IQ2_XS, 128K context, 8-bit KV cache, MTP on, KV streaming on, images on, speed projection off. The IQ3_XXS tier ran at 98.1 tok/s whole-request with ~120 tok/s peak decode and 131 tok/s prefill; the user notes it was 3.7x longer and qualitatively deeper for only 8% less speed. Both tiers match the README's 24 GB card estimate of 100-140 t/s.