llamaperf
Sep 29, 2026
Throughput
26.1 t/s gen · 310.0 t/s pp
Quant
IQ2_M (GGUF)
KV cache
q8_0
System RAM
128 GB

Summary

User reports MiMo 2.6 Flash-RL at 26.1 tok/s decode (tg128) and ~310 tok/s prefill (pp4096) on a Strix Halo 128GB machine. Setup is llama.cpp (Vulkan, Radeon 8060S) with a custom IQ2_M-class GGUF at 2.76 bpw and q8_0 KV cache, 32768 context, -ub 2048. Prefill drops to 193 tok/s at the default -ub 512. The 100.4 GiB quant is measured against the native MXFP4 GGUF: KLD 0.164 mean, 88.2% same top-1 token, PPL ratio 1.128. MTP self-speculation is included in the file but does not speed up decode (draft acceptance ~50%), so speculation is off.