agenticcodinglong-context
Mimo 2.5 uses 5-to-1 sliding-window attention similar to Gemma 3, stays fast at large context on dual RTX Pro 6000. Also mentions Step 3.7 Flash (3-to-1) at ~40 t/s at 178k context. MiniMax M3 and DeepSeek V4 have kernel issues on Blackwell consumer GPUs.