llamaperf

Mimo 2.6 309B (15B active) Flash-RL

on 2× NVIDIA DGX Spark · vLLM · 300,000 ctx

Tone: mixed
Oct 3, 2026
Throughput
53.3 t/s gen
Quant
MXFP4 (MXFP4)
KV cache
fp8

Use cases

codingagenticvisionlong-contexttool-use

Summary

User reports MiMo-V2.6-Flash-RL at 53.31 tok/s per stream at C1 on two DGX Sparks (GB10, 121.7 GiB unified memory each) with vLLM tensor parallel 2 and DFlash speculative decoding (7 draft tokens). Setup is vLLM with MXFP4 experts, fp8 KV cache, 300K max context, marlin MoE backend, GPU memory utilization 0.90, max-num-seqs 8, KV pool 1,835,052 tokens. Aggregate throughput is 155.77 tok/s at six streams; cold prefill ranges from 1,946.6 tok/s at 2K to 656.4 tok/s at 248K; TTFT 0.367 s at C1. User notes tool-call storms in agent use and that prose/narrative decode is slow due to low DFlash acceptance.