llamaperf

Mimo 2.6 309B (15B active) Flash

on 2× NVIDIA DGX Spark · vLLM · 1,048,576 ctx

Tone: mixed
Sep 22, 2026
Throughput
42-51 t/s gen
KV cache
fp8

Use cases

agentictool-usecodinglong-context

Summary

User reports MiMo-V2.6-Flash-RL at ~42–51 t/s on 2× DGX Spark (GB10) with vLLM, TP=2 over RoCE, fp8 KV cache, full 1,048,576 context, and DFlash with 7 draft tokens. Setup uses the tonyd2wild recipe's sm121-v11-dflash2 image. Decode is ~42–51 t/s on agentic code at default sampling, 69 t/s on the recipe's coding benchmark at temperature 0, and ~20–23 t/s on prose. Deep prefill is the weak spot at ~160 t/s past 800K, so a cold 1M prompt takes ~67 min. User documents three serving bugs: empty streaming replies with thinking on (fixed by pre-opening the thinking tag in the chat template), dropped earlier reasoning in tool loops (fixed by template and chat_utils.py fallbacks), and a hidden 2,048-token output cap from generation_config.json (fixed with --override-generation-config max_new_tokens 131072). A 1M needle test retrieved 3/3 hidden codes from a 996K-token prompt.