llamaperf

Mimo 2.6 309B (15B active) Flash-RL

on 5× NVIDIA RTX 5090 · mimo26f-afd · 1,048,576 ctx

Oct 3, 2026
Throughput
109.7 t/s gen · 5123.0 t/s pp
Quant
MXFP4 (MXFP4)
KV cache
FP8

Use cases

visiontool-uselong-contextagentic

Summary

User reports MiMo-V2.6-Flash-RL at 109.7 tok/s decode on one stream on one RTX 5090 plus four DGX Spark (GB10) systems. Setup is the mimo26f-afd engine with MXFP4 weights and FP8 KV cache, attention-FFN disaggregation over RoCE v2 RDMA, DFlash speculative decoding, up to 1,048,576 tokens of context. Decode reaches 279.1 tok/s across six streams and 416.7 tok/s across sixteen; cold prefill is 3,630 / 5,123 / 4,816 / 4,283 tok/s at 2K / 8K / 32K / 64K. The comparison is a 4-Spark vLLM TP4 reference without the 5090, so the uplift includes the fifth device.