llamaperf
Sep 29, 2026
Throughput
246.0 t/s gen
KV cache
FP8

Use cases

visionmultilinguallong-context

Summary

User reports MiMo-V2.6-Flash-RL at 246 tok/s decode at C1 (90.0 decode steps/s) on two RTX PRO 6000 Blackwell GPUs. Setup is vLLM with FP8 KV cache, TP2 at 0.985 utilization, DFlash drafter with 7 tokens, video off, 1.31M token KV cache. Decode is 42.3 steps/s at C8 and 30.6 at C16; 246 tok/s at C1 on mixed real tasks. Prefill is 9.7K / 9.1K tok/s at 8K / 32K. First start autotunes b12x for about 15 minutes.