llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Mimo 2.6
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Mimo 2.6 Flash-MOPD

AMD Strix Halo 128GB · llama.cpp · 262,144 ctx

Tone: mixed
reported speed:
23.6 tokens/s generation · 37.5 tokens/s prompt processing
quant:
MQ-IQ2-XXS-XS-Q8 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcodingtool-use

User reports MiMo V2.6 Flash MOPD mixed GGUF at 23.59 t/s decode and 37.54 t/s prefill on an AMD Strix Halo 128GB. Setup is llama.cpp (Vulkan) with MQ-IQ2-XXS-XS-Q8 GGUF and q8_0 KV cache, 256K context, thinking disabled, one request at a time. Longer inputs slow decode to 15.89 t/s at 4,860 tokens and 7.39 t/s at 48,600 tokens. MTP was disabled; a preliminary test showed MTP 3 taking 225.591 s versus 39.896 s without it. An agent benchmark scored 330/900 (36.7%) with all ten tasks timing out.

Oct 7, 2026
reported speed:
49.6 tokens/s generation · 2370.2 tokens/s prompt processing
quant:
2.20 bpw (EXL3)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo-V2.6-Flash-RL at 49.57 tok/s decode (p50, single stream) on an NVIDIA RTX 6000 Ada 96GB. Setup is ExLlamaV3 (vcruz305 fork) with a 2.20 bpw EXL3 pack (86.94 GB) at 65,536-token context and a 65,536-token KV pool; prefill measured 2,370.2 tok/s p50. With the DFlash drafter attached, decode rises to 184.11 tok/s p50 and prefill is 2,257.7 tok/s. Aggregate throughput under load reaches 236.2 tok/s at C=8 without the drafter and 330.6 tok/s with it; the eight-stream per-stream decode p50 is 34.95 and 54.46 tok/s respectively. The 2.50 bpw pack (98.48 GB) does not fit a 96 GB card. On a DGX Spark / GB10 unified-memory host the 2.50 bpw pack runs at 34.9 tok/s decode (code) and 23.7 tok/s (prose) with draft_accept 0.78.

Oct 6, 2026
Tone: mixed
reported speed:
53.3 tokens/s generation
quant:
MXFP4 (MXFP4)
kv:
fp8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticvisionlong-contexttool-use

User reports MiMo-V2.6-Flash-RL at 53.31 tok/s per stream at C1 on two DGX Sparks (GB10, 121.7 GiB unified memory each) with vLLM tensor parallel 2 and DFlash speculative decoding (7 draft tokens). Setup is vLLM with MXFP4 experts, fp8 KV cache, 300K max context, marlin MoE backend, GPU memory utilization 0.90, max-num-seqs 8, KV pool 1,835,052 tokens. Aggregate throughput is 155.77 tok/s at six streams; cold prefill ranges from 1,946.6 tok/s at 2K to 656.4 tok/s at 248K; TTFT 0.367 s at C1. User notes tool-call storms in agent use and that prose/narrative decode is slow due to low DFlash acceptance.

Oct 3, 2026

Mimo 2.6 309B (15B active) Flash-RL

5× NVIDIA RTX 5090 · mimo26f-afd · 1,048,576 ctx

reported speed:
109.7 tokens/s generation · 5123.0 tokens/s prompt processing
quant:
MXFP4 (MXFP4)
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visiontool-uselong-contextagentic

User reports MiMo-V2.6-Flash-RL at 109.7 tok/s decode on one stream on one RTX 5090 plus four DGX Spark (GB10) systems. Setup is the mimo26f-afd engine with MXFP4 weights and FP8 KV cache, attention-FFN disaggregation over RoCE v2 RDMA, DFlash speculative decoding, up to 1,048,576 tokens of context. Decode reaches 279.1 tok/s across six streams and 416.7 tok/s across sixteen; cold prefill is 3,630 / 5,123 / 4,816 / 4,283 tok/s at 2K / 8K / 32K / 64K. The comparison is a 4-Spark vLLM TP4 reference without the 5090, so the uplift includes the fifth device.

Oct 3, 2026
reported speed:
246.0 tokens/s generation
kv:
FP8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

visionmultilinguallong-context

User reports MiMo-V2.6-Flash-RL at 246 tok/s decode at C1 (90.0 decode steps/s) on two RTX PRO 6000 Blackwell GPUs. Setup is vLLM with FP8 KV cache, TP2 at 0.985 utilization, DFlash drafter with 7 tokens, video off, 1.31M token KV cache. Decode is 42.3 steps/s at C8 and 30.6 at C16; 246 tok/s at C1 on mixed real tasks. Prefill is 9.7K / 9.1K tok/s at 8K / 32K. First start autotunes b12x for about 15 minutes.

Sep 29, 2026
reported speed:
26.1 tokens/s generation · 310.0 tokens/s prompt processing
quant:
IQ2_M (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo 2.6 Flash-RL at 26.1 tok/s decode (tg128) and ~310 tok/s prefill (pp4096) on a Strix Halo 128GB machine. Setup is llama.cpp (Vulkan, Radeon 8060S) with a custom IQ2_M-class GGUF at 2.76 bpw and q8_0 KV cache, 32768 context, -ub 2048. Prefill drops to 193 tok/s at the default -ub 512. The 100.4 GiB quant is measured against the native MXFP4 GGUF: KLD 0.164 mean, 88.2% same top-1 token, PPL ratio 1.128. MTP self-speculation is included in the file but does not speed up decode (draft acceptance ~50%), so speculation is off.

Sep 29, 2026
reported speed:
41.6 tokens/s generation · 1718.3 tokens/s prompt processing
quant:
Q8_0 (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports MiMo 2.6 Distill Qwen 9B at 41.63 t/s generation and 1718.29 t/s prompt processing on an RTX 5060 Ti 16GB. Setup is llama.cpp (Llama UI) with GGUF Q8_0 weights, q4_0 KV cache, 122880-token context, and Flash Attention enabled. The run produced 11979 output tokens over 4 min 47 s from a 1929-token prompt; the first generated version failed and a second review pass by the model produced the final working demo.

Sep 27, 2026
reported speed:
68.3 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports MiMo V2.6 Pro RL at 68.3 t/s on eight DGX Spark units, up from 17.8 t/s without speculation. Setup is vLLM with DFlash speculative decoding on official weights; the figure includes prefill and request overhead. Outputs matched byte for byte across 12 cases, including two baseline mistakes. Four concurrent requests and a 257K-token input were also tested.

Sep 23, 2026
Tone: mixed
reported speed:
49.1 tokens/s generation · 1238.4 tokens/s prompt processing
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Mimo 2.6 Flash at 49.1 t/s generation and 1238.4 t/s prompt processing on an M5 Ultra 256GB, averaged over five trials at 32,000 prompt tokens and 64 generation tokens. Setup is a custom mlx-vlm patch loading the original weights, with MTP enabled but apparently not working. A second run at 64,000 prompt tokens and 1024 generation tokens averaged 38.1 t/s generation and 981.9 t/s prompt processing. User also compares Qwen3.8 Flash-Next FP8 on oMLX, reaching 39.7 to 43.9 t/s generation across 32,768 to 200,000 token prompts.

Sep 22, 2026
Showing 1–9 of 9
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23