llamaperf
Oct 5, 2026
Throughput
30.5 t/s gen · 250.4 t/s pp
Quant
IQ1_M (GGUF)
KV cache
q8_0
System RAM
32 GB

Use cases

coding

Summary

User reports Qwen3.8-Flash-Next at 30.5 t/s generation and 250.4 t/s prompt processing on 2x RTX 5060 Ti with 32GB system RAM. Setup is llama.cpp with an IQ1_M GGUF (27.58 GB), 98304 context, q8_0 KV cache, tensor split 1,1, and 8 MoE layers offloaded to CPU. User asks whether their settings are correct and what config others run for this model on this hardware.