llamaperf

Minimax M2.1

1 report

As of 7 Oct 2026, Minimax M2.1 230B · 10B active at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

Minimax M2.1 VRAM requirements by size and quant →

How does Minimax M2.1 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Minimax M2.1 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Minimax M2.1 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Minimax M2.1

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: negative
reported speed:
25.0 tokens/s generation
quant:
4bit

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcodinglong-context

User reports Minimax-m2.1-4bit at 25 t/s generation on an M3 Ultra 256GB with 80-core GPU, compared against llama.cpp with flash attention at 32 t/s. Setup is MLX with 4bit weights at 30K context; the same model under llama.cpp with flash attention reached 32 t/s generation and 78s prompt processing. At 146K context MLX fell to 5.95 t/s generation versus 12.12 t/s for llama.cpp, and the user says MLX runs about 50% slower on long-context generation, impacting agentic coding.

Oct 4, 2026

Get a weekly email of new Minimax M2.1 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Minimax M2.1 230B (10B active)
M3 Ultra 256GB
4bit
MLX
30,00025.0 tokens/s