Mimo 2.6 Flash-MOPD
AMD Strix Halo 128GB · llama.cpp · 262,144 ctx
- reported speed:
- 23.6 tokens/s generation · 37.5 tokens/s prompt processing
- quant:
- MQ-IQ2-XXS-XS-Q8 (GGUF)
- kv:
- q8_0
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports MiMo V2.6 Flash MOPD mixed GGUF at 23.59 t/s decode and 37.54 t/s prefill on an AMD Strix Halo 128GB. Setup is llama.cpp (Vulkan) with MQ-IQ2-XXS-XS-Q8 GGUF and q8_0 KV cache, 256K context, thinking disabled, one request at a time. Longer inputs slow decode to 15.89 t/s at 4,860 tokens and 7.39 t/s at 48,600 tokens. MTP was disabled; a preliminary test showed MTP 3 taking 225.591 s versus 39.896 s without it. An agent benchmark scored 330/900 (36.7%) with all ten tasks timing out.