llamaperf

MiniMax M2.7

1 report

As of 11 Oct 2026, MiniMax M2.7 230B · 10B active at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

MiniMax M2.7 VRAM requirements by size and quant →

How does MiniMax M2.7 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for MiniMax M2.7 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run MiniMax M2.7 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for MiniMax M2.7

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
Apple hardware
generation:
32.7 tokens/s
prompt processing (prefill):
420.8 tokens/s
quant:
UD-Q4_K_M (GGUF)
kv:
Q8
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.

User reports MiniMax M2.7 at 32.7 t/s generation and 420.8 t/s prefill on an M2 Ultra 192GB. Setup is llama.cpp at UD-Q4_K_M, one request, speculative decoding off. Token-weighted rates across four sequential agentic tasks in crypdick's run of 2026-08-10. A source-reported measurement, not independently reproduced.

Oct 10, 2026

Get a weekly email of new MiniMax M2.7 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
MiniMax M2.7 230B (10B active)
M2 Ultra 192GB
UD-Q4_K_M
llama.cpp
Not reported32.7 tokens/s