Oct 3, 2026
Summary
User reports 850 t/s prefill and 50 t/s generation with MTP at temperature 0 on an AMD RX 9060 XT 16GB, running Qwen3.8-27B-GSQ-RCO IQ3_S in a custom inference engine.
MTP is used for the generation figure; the user notes MTP with temperature above 0 is not developed yet.
User asks for ideas to systematically test the engine, having already run KL divergence against BF12 on CPU, bit correctness checks, needle-in-a-haystack at 25/50/75/90% key positions with distractors (also against llama.cpp), and HumanEval (92/93 passed so far).