Mimo 2.6 309B (15B active) Flash-RL
on 5× NVIDIA RTX 5090 · mimo26f-afd · 1,048,576 ctx
Oct 3, 2026
Use cases
visiontool-uselong-contextagentic
Summary
User reports MiMo-V2.6-Flash-RL at 109.7 tok/s decode on one stream on one RTX 5090 plus four DGX Spark (GB10) systems.
Setup is the mimo26f-afd engine with MXFP4 weights and FP8 KV cache, attention-FFN disaggregation over RoCE v2 RDMA, DFlash speculative decoding, up to 1,048,576 tokens of context.
Decode reaches 279.1 tok/s across six streams and 416.7 tok/s across sixteen; cold prefill is 3,630 / 5,123 / 4,816 / 4,283 tok/s at 2K / 8K / 32K / 64K. The comparison is a 4-Spark vLLM TP4 reference without the 5090, so the uplift includes the fifth device.