llamaperf

Mimo 2.6 309B (15B active) Flash

on 4× Unknown GPU · 262,144 ctx

Oct 3, 2026
Throughput
5.6 t/s gen

Summary

User reports MiMo-V2.6-Flash running a 262144-token prompt end to end on four ranks, with a decode step at 8 tokens of context going from 3.63 to 5.63 tokens a second. Setup is a four-rank run of the model's own code, with a 256k prompt reaching 48.32 tok/s and a 64k prompt at chunk 4096 reaching 104.4 tok/s (95.5 at chunk 2048). The gains come from a fixed 2^26-score row step in the attention block loop and from cutting 145 cudaStreamSynchronize calls a decode step; a free-memory budget instead of the constant measured 8246.7 s at 256k against 5424.9 s for the constant.