llamaperf
Oct 4, 2026
Throughput
80-110 t/s gen
Quant
Q4_K_XL (GGUF)
System RAM
96 GB

Summary

User reports Qwen3.8 Flash-Next at 80-110 tok/s on 2x RTX 3090 with 96GB DDR5 system RAM. Setup uses the Strata engine with a Q4_K_XL GGUF quant from bitlamas, at 2500 prompt tokens. The user calls it the biggest thing since Qwen 3.8 27B, offering the speed of Qwen 3.5 35B-A3B with higher intelligence than 3.8 27B.