Mimo 2.6 Flash-MOPD
on AMD Strix Halo 128GB · llama.cpp · 262,144 ctx
Oct 7, 2026
Use cases
agenticcodingtool-use
Summary
User reports MiMo V2.6 Flash MOPD mixed GGUF at 23.59 t/s decode and 37.54 t/s prefill on an AMD Strix Halo 128GB.
Setup is llama.cpp (Vulkan) with MQ-IQ2-XXS-XS-Q8 GGUF and q8_0 KV cache, 256K context, thinking disabled, one request at a time.
Longer inputs slow decode to 15.89 t/s at 4,860 tokens and 7.39 t/s at 48,600 tokens. MTP was disabled; a preliminary test showed MTP 3 taking 225.591 s versus 39.896 s without it. An agent benchmark scored 330/900 (36.7%) with all ten tasks timing out.