llamaperf

Mimo 2.6 Flash-MOPD

on AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: mixed
Oct 7, 2026
Throughput
23.6 t/s gen · 37.5 t/s pp
Quant
MQ-IQ2-XXS-XS-Q8 (GGUF)
KV cache
q8_0
System RAM
128 GB

Use cases

agenticcodingtool-use

Summary

User reports MiMo V2.6 Flash MOPD mixed GGUF at 23.59 t/s decode and 37.54 t/s prefill on an AMD Strix Halo 128GB. Setup is llama.cpp (Vulkan) with MQ-IQ2-XXS-XS-Q8 GGUF and q8_0 KV cache, 256K context, thinking disabled, one request at a time. Longer inputs slow decode to 15.89 t/s at 4,860 tokens and 7.39 t/s at 48,600 tokens. MTP was disabled; a preliminary test showed MTP 3 taking 225.591 s versus 39.896 s without it. An agent benchmark scored 330/900 (36.7%) with all ten tasks timing out.