llamaperf
Sep 27, 2026
Throughput
41.6 t/s gen · 1718.3 t/s pp
Quant
Q8_0 (GGUF)
KV cache
q4_0
VRAM reported
16 GB

Use cases

coding

Summary

User reports MiMo 2.6 Distill Qwen 9B at 41.63 t/s generation and 1718.29 t/s prompt processing on an RTX 5060 Ti 16GB. Setup is llama.cpp (Llama UI) with GGUF Q8_0 weights, q4_0 KV cache, 122880-token context, and Flash Attention enabled. The run produced 11979 output tokens over 4 min 47 s from a 1929-token prompt; the first generated version failed and a second review pass by the model produced the final working demo.