llamaperf
Oct 6, 2026
Throughput
49.6 t/s gen · 2370.2 t/s pp
Quant
2.20 bpw (EXL3)
VRAM reported
96 GB

Summary

User reports MiMo-V2.6-Flash-RL at 49.57 tok/s decode (p50, single stream) on an NVIDIA RTX 6000 Ada 96GB. Setup is ExLlamaV3 (vcruz305 fork) with a 2.20 bpw EXL3 pack (86.94 GB) at 65,536-token context and a 65,536-token KV pool; prefill measured 2,370.2 tok/s p50. With the DFlash drafter attached, decode rises to 184.11 tok/s p50 and prefill is 2,257.7 tok/s. Aggregate throughput under load reaches 236.2 tok/s at C=8 without the drafter and 330.6 tok/s with it; the eight-stream per-stream decode p50 is 34.95 and 54.46 tok/s respectively. The 2.50 bpw pack (98.48 GB) does not fit a 96 GB card. On a DGX Spark / GB10 unified-memory host the 2.50 bpw pack runs at 34.9 tok/s decode (code) and 23.7 tok/s (prose) with draft_accept 0.78.