llamaperf
Oct 3, 2026
Throughput
50.0 t/s gen · 850.0 t/s pp
Quant
IQ3_S (GGUF)
VRAM reported
16 GB

Summary

User reports 850 t/s prefill and 50 t/s generation with MTP at temperature 0 on an AMD RX 9060 XT 16GB, running Qwen3.8-27B-GSQ-RCO IQ3_S in a custom inference engine. MTP is used for the generation figure; the user notes MTP with temperature above 0 is not developed yet. User asks for ideas to systematically test the engine, having already run KL divergence against BF12 on CPU, bit correctness checks, needle-in-a-haystack at 25/50/75/90% key positions with distractors (also against llama.cpp), and HumanEval (92/93 passed so far).