Minimax M2.1 230B (10B active)
M3 Ultra 256GB · MLX · 30,000 ctx
- reported speed:
- 25.0 tokens/s generation
- quant:
- 4bit
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Minimax-m2.1-4bit at 25 t/s generation on an M3 Ultra 256GB with 80-core GPU, compared against llama.cpp with flash attention at 32 t/s. Setup is MLX with 4bit weights at 30K context; the same model under llama.cpp with flash attention reached 32 t/s generation and 78s prompt processing. At 146K context MLX fell to 5.95 t/s generation versus 12.12 t/s for llama.cpp, and the user says MLX runs about 50% slower on long-context generation, impacting agentic coding.