MiniMax M2.7 230B (10B active)
M2 Ultra 192GB · llama.cpp
- generation:
- 32.7 tokens/s
- prompt processing (prefill):
- 420.8 tokens/s
- quant:
- UD-Q4_K_M (GGUF)
- kv:
- Q8
- flash attention:
- on
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.
User reports MiniMax M2.7 at 32.7 t/s generation and 420.8 t/s prefill on an M2 Ultra 192GB. Setup is llama.cpp at UD-Q4_K_M, one request, speculative decoding off. Token-weighted rates across four sequential agentic tasks in crypdick's run of 2026-08-10. A source-reported measurement, not independently reproduced.