Minimax M3
2× M3 Ultra 256GB · oMLX
- reported speed:
- 17.3 tokens/s generation
- quant:
- 4bit (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
Post describes a bug in oMLX 0.6.4 distributed clustering where the coordinator (M3 Ultra 256GB) fails to release RAM/GPU after a crash. Model: mlx-community/MiniMax-M3-4bit (236GB). Cluster: rank 0 Mac Studio M3 Ultra 256GB, rank 1 Mac Studio M2 Ultra 192GB. First completion: 17 tokens, prompt 7,693, 17.3 tok/s. Crash triggered by overlapping requests with pipeline_prefill_overlap and coalesced batching. Post-crash: ~116GB wired memory with no owning process, GPU pinned at 100%, requires reboot. Worker (M2 Ultra) released memory cleanly. User also mentions MiniMax-M3 crashes the cluster after first prompt. No subjective rating given.