llamaperf

MiniMax M2.7 230B (10B active)

on M2 Ultra 192GB · llama.cpp

Apple hardware
Oct 10, 2026
Throughput
32.7 t/s gen · 420.8 tokens/s prompt processing (prefill)Prompt processing measures input prefill speed; it does not tell you the time to first token.
Quant
UD-Q4_K_M (GGUF)
KV cache
Q8
Flash Attention
on
System RAM
192 GB

Summary

User reports MiniMax M2.7 at 32.7 t/s generation and 420.8 t/s prefill on an M2 Ultra 192GB. Setup is llama.cpp at UD-Q4_K_M, one request, speculative decoding off. Token-weighted rates across four sequential agentic tasks in crypdick's run of 2026-08-10. A source-reported measurement, not independently reproduced.