llamaperf
Oct 3, 2026
Throughput
167.0 t/s gen · 2466.0 t/s pp
Quant
IQ3_S (GGUF)

Summary

User reports Qwen3.8 Flash-Next at 2466 t/s prompt processing and 167 t/s generation on an RTX 3090 plus RTX 5070 Ti with the Strata engine and IQ3_S quant. Setup is Strata with IQ3_S; the user also added UD-Q4_K_XL support, which reached 2341 t/s prompt and 126 t/s generation. The user doubled Strata throughput over a week of profiling and reached 5x llama.cpp on UD-Q4_K_XL, where the starting point was 6 t/s. Earlier steps were 21 t/s, 27 t/s and 51 t/s on IQ3_XXS.