llamaperf

Minimax M3

MiniMax · 1 report

Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: negative
reported speed:
17.3 tokens/s generation
quant:
4bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post describes a bug in oMLX 0.6.4 distributed clustering where the coordinator (M3 Ultra 256GB) fails to release RAM/GPU after a crash. Model: mlx-community/MiniMax-M3-4bit (236GB). Cluster: rank 0 Mac Studio M3 Ultra 256GB, rank 1 Mac Studio M2 Ultra 192GB. First completion: 17 tokens, prompt 7,693, 17.3 tok/s. Crash triggered by overlapping requests with pipeline_prefill_overlap and coalesced batching. Post-crash: ~116GB wired memory with no owning process, GPU pinned at 100%, requires reboot. Worker (M2 Ultra) released memory cleanly. User also mentions MiniMax-M3 crashes the cluster after first prompt. No subjective rating given.