llamaperf
Oct 5, 2026
Throughput
120.1 t/s gen · 4700-4800 t/s pp
Quant
IQ3_S (GGUF)
KV cache
int8
System RAM
192 GB
VRAM reported
24 GB

Use cases

codingagenticvisiontool-use

Summary

User reports Qwen3.8-Flash-Next at 120.10 tok/s decode at 32K context on one RTX 4090, with 112.55 tok/s at 128K and roughly 4,700-4,800 tok/s on uncached 32K prefill. Setup is Strata with GSQ IQ3_S GGUF weights, int8 KV cache and a 32K resident window, MTP speculation with four draft tokens, 11 pool workers and a 0.35 PCIe fraction; all experts and the IQ4_NL ngram table stay in RAM, so the run is partially CPU-offloaded. The host is a Ryzen 7900 with 192 GB RAM at a 280 W GPU limit. The 120.10 tok/s figure is the median of launch medians for the selected 24-expert-cache profile, a +0.88% change over the 96-swap baseline; the 128K figure is +4.26%. Earlier matched decode cells give 110.8 tok/s median for the text profile and 97.65 tok/s for the GPU vision profile, and a matched A/B isolates GPU vision residency at 114.85 versus 104.25 tok/s (-9.2%). The 397B partial-offload probe reached 12.42 tok/s with quality unqualified.