llamaperf

Mimo 2.6

Xiaomi · 11 reports

As of 7 Oct 2026, Mimo 2.6 309B · 15B active at 2-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

Mimo 2.6 VRAM requirements by size and quant →

How does Mimo 2.6 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Mimo 2.6 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Mimo 2.6 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Mimo 2.6

Filter this model’s reports by setup →

Mimo 2.6 Flash-MOPD

AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: mixed
reported speed:
23.6 tokens/s generation · 37.5 tokens/s prompt processing
quant:
MQ-IQ2-XXS-XS-Q8 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcodingtool-use

User reports MiMo V2.6 Flash MOPD mixed GGUF at 23.59 t/s decode and 37.54 t/s prefill on an AMD Strix Halo 128GB. Setup is llama.cpp (Vulkan) with MQ-IQ2-XXS-XS-Q8 GGUF and q8_0 KV cache, 256K context, thinking disabled, one request at a time. Longer inputs slow decode to 15.89 t/s at 4,860 tokens and 7.39 t/s at 48,600 tokens. MTP was disabled; a preliminary test showed MTP 3 taking 225.591 s versus 39.896 s without it. An agent benchmark scored 330/900 (36.7%) with all ten tasks timing out.

Oct 7, 2026
reported speed:
49.6 tokens/s generation · 2370.2 tokens/s prompt processing
quant:
2.20 bpw (EXL3)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo-V2.6-Flash-RL at 49.57 tok/s decode (p50, single stream) on an NVIDIA RTX 6000 Ada 96GB. Setup is ExLlamaV3 (vcruz305 fork) with a 2.20 bpw EXL3 pack (86.94 GB) at 65,536-token context and a 65,536-token KV pool; prefill measured 2,370.2 tok/s p50. With the DFlash drafter attached, decode rises to 184.11 tok/s p50 and prefill is 2,257.7 tok/s. Aggregate throughput under load reaches 236.2 tok/s at C=8 without the drafter and 330.6 tok/s with it; the eight-stream per-stream decode p50 is 34.95 and 54.46 tok/s respectively. The 2.50 bpw pack (98.48 GB) does not fit a 96 GB card. On a DGX Spark / GB10 unified-memory host the 2.50 bpw pack runs at 34.9 tok/s decode (code) and 23.7 tok/s (prose) with draft_accept 0.78.

Oct 6, 2026

Mimo 2.6 309B (15B active) Flash

4× Unknown GPU · 262,144 ctx

reported speed:
5.6 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo-V2.6-Flash running a 262144-token prompt end to end on four ranks, with a decode step at 8 tokens of context going from 3.63 to 5.63 tokens a second. Setup is a four-rank run of the model's own code, with a 256k prompt reaching 48.32 tok/s and a 64k prompt at chunk 4096 reaching 104.4 tok/s (95.5 at chunk 2048). The gains come from a fixed 2^26-score row step in the attention block loop and from cutting 145 cudaStreamSynchronize calls a decode step; a free-memory budget instead of the constant measured 8246.7 s at 256k against 5424.9 s for the constant.

Oct 3, 2026
Tone: mixed
reported speed:
53.3 tokens/s generation
quant:
MXFP4 (MXFP4)
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticvisionlong-contexttool-use

User reports MiMo-V2.6-Flash-RL at 53.31 tok/s per stream at C1 on two DGX Sparks (GB10, 121.7 GiB unified memory each) with vLLM tensor parallel 2 and DFlash speculative decoding (7 draft tokens). Setup is vLLM with MXFP4 experts, fp8 KV cache, 300K max context, marlin MoE backend, GPU memory utilization 0.90, max-num-seqs 8, KV pool 1,835,052 tokens. Aggregate throughput is 155.77 tok/s at six streams; cold prefill ranges from 1,946.6 tok/s at 2K to 656.4 tok/s at 248K; TTFT 0.367 s at C1. User notes tool-call storms in agent use and that prose/narrative decode is slow due to low DFlash acceptance.

Oct 3, 2026

Mimo 2.6 309B (15B active) Flash-RL

5× NVIDIA RTX 5090 · mimo26f-afd · 1,048,576 ctx

reported speed:
109.7 tokens/s generation · 5123.0 tokens/s prompt processing
quant:
MXFP4 (MXFP4)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visiontool-uselong-contextagentic

User reports MiMo-V2.6-Flash-RL at 109.7 tok/s decode on one stream on one RTX 5090 plus four DGX Spark (GB10) systems. Setup is the mimo26f-afd engine with MXFP4 weights and FP8 KV cache, attention-FFN disaggregation over RoCE v2 RDMA, DFlash speculative decoding, up to 1,048,576 tokens of context. Decode reaches 279.1 tok/s across six streams and 416.7 tok/s across sixteen; cold prefill is 3,630 / 5,123 / 4,816 / 4,283 tok/s at 2K / 8K / 32K / 64K. The comparison is a 4-Spark vLLM TP4 reference without the 5090, so the uplift includes the fifth device.

Oct 3, 2026
reported speed:
246.0 tokens/s generation
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionmultilinguallong-context

User reports MiMo-V2.6-Flash-RL at 246 tok/s decode at C1 (90.0 decode steps/s) on two RTX PRO 6000 Blackwell GPUs. Setup is vLLM with FP8 KV cache, TP2 at 0.985 utilization, DFlash drafter with 7 tokens, video off, 1.31M token KV cache. Decode is 42.3 steps/s at C8 and 30.6 at C16; 246 tok/s at C1 on mixed real tasks. Prefill is 9.7K / 9.1K tok/s at 8K / 32K. First start autotunes b12x for about 15 minutes.

Sep 29, 2026
reported speed:
26.1 tokens/s generation · 310.0 tokens/s prompt processing
quant:
IQ2_M (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo 2.6 Flash-RL at 26.1 tok/s decode (tg128) and ~310 tok/s prefill (pp4096) on a Strix Halo 128GB machine. Setup is llama.cpp (Vulkan, Radeon 8060S) with a custom IQ2_M-class GGUF at 2.76 bpw and q8_0 KV cache, 32768 context, -ub 2048. Prefill drops to 193 tok/s at the default -ub 512. The 100.4 GiB quant is measured against the native MXFP4 GGUF: KLD 0.164 mean, 88.2% same top-1 token, PPL ratio 1.128. MTP self-speculation is included in the file but does not speed up decode (draft acceptance ~50%), so speculation is off.

Sep 29, 2026
reported speed:
41.6 tokens/s generation · 1718.3 tokens/s prompt processing
quant:
Q8_0 (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports MiMo 2.6 Distill Qwen 9B at 41.63 t/s generation and 1718.29 t/s prompt processing on an RTX 5060 Ti 16GB. Setup is llama.cpp (Llama UI) with GGUF Q8_0 weights, q4_0 KV cache, 122880-token context, and Flash Attention enabled. The run produced 11979 output tokens over 4 min 47 s from a 1929-token prompt; the first generated version failed and a second review pass by the model produced the final working demo.

Sep 27, 2026
reported speed:
68.3 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo V2.6 Pro RL at 68.3 t/s on eight DGX Spark units, up from 17.8 t/s without speculation. Setup is vLLM with DFlash speculative decoding on official weights; the figure includes prefill and request overhead. Outputs matched byte for byte across 12 cases, including two baseline mistakes. Four concurrent requests and a 257K-token input were also tested.

Sep 23, 2026
Tone: mixed
reported speed:
42-51 tokens/s generation
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentictool-usecodinglong-context

User reports MiMo-V2.6-Flash-RL at ~42–51 t/s on 2× DGX Spark (GB10) with vLLM, TP=2 over RoCE, fp8 KV cache, full 1,048,576 context, and DFlash with 7 draft tokens. Setup uses the tonyd2wild recipe's sm121-v11-dflash2 image. Decode is ~42–51 t/s on agentic code at default sampling, 69 t/s on the recipe's coding benchmark at temperature 0, and ~20–23 t/s on prose. Deep prefill is the weak spot at ~160 t/s past 800K, so a cold 1M prompt takes ~67 min. User documents three serving bugs: empty streaming replies with thinking on (fixed by pre-opening the thinking tag in the chat template), dropped earlier reasoning in tool loops (fixed by template and chat_utils.py fallbacks), and a hidden 2,048-token output cap from generation_config.json (fixed with --override-generation-config max_new_tokens 131072). A 1M needle test retrieved 3/3 hidden codes from a 996K-token prompt.

Sep 22, 2026
Tone: mixed
reported speed:
49.1 tokens/s generation · 1238.4 tokens/s prompt processing
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Mimo 2.6 Flash at 49.1 t/s generation and 1238.4 t/s prompt processing on an M5 Ultra 256GB, averaged over five trials at 32,000 prompt tokens and 64 generation tokens. Setup is a custom mlx-vlm patch loading the original weights, with MTP enabled but apparently not working. A second run at 64,000 prompt tokens and 1024 generation tokens averaged 38.1 t/s generation and 981.9 t/s prompt processing. User also compares Qwen3.8 Flash-Next FP8 on oMLX, reaching 39.7 to 43.9 t/s generation across 32,768 to 200,000 token prompts.

Sep 22, 2026
Engines people run Mimo 2.6 with
EngineReports
vLLM4
llama.cpp3
ExLlamaV31
mlx-vlm1

Get a weekly email of new Mimo 2.6 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Mimo 2.6 Flash-MOPD
AMD Strix Halo 128GB
MQ-IQ2-XXS-XS-Q8
llama.cpp
262,14423.6 tokens/s
Mimo 2.6 309B (15B active) Flash-RL
NVIDIA RTX 6000 Ada
2.20 bpw
ExLlamaV3
65,53649.6 tokens/s
Mimo 2.6 309B (15B active) Flash-RL
2× NVIDIA DGX Spark
MXFP4
vLLM
300,00053.3 tokens/s
Mimo 2.6 309B (15B active) Flash-RL
5× NVIDIA RTX 5090
MXFP4
mimo26f-afd
1,048,576109.7 tokens/s
Mimo 2.6 309B (15B active) Flash-RL
2× NVIDIA RTX Pro 6000 Blackwell
Not reported
vLLM
1,000,000246.0 tokens/s
Mimo 2.6 309B (15B active) Flash-RL
AMD Strix Halo 128GB
IQ2_M
llama.cpp
32,76826.1 tokens/s