llamaperf
Oct 5, 2026
Throughput
34.0 t/s gen · 520.0 t/s pp
Quant
Q8 (GGUF)
System RAM
256 GB

Summary

User reports Qwen3.8 Flash-Next at Q8 running at 34 t/s generation and up to 520 t/s prefill on 2x RTX 3090 with 256GB of system RAM. Setup uses the Strata engine with MTP set to 5; most of the model sits in system RAM rather than GPU memory. For comparison, the user says llama.cpp without MTP gives 13 t/s generation and 115 t/s prefill on the same setup, and notes MTP usually does not help when most of the model is in system RAM.