mlx-serve
An inference engine for running open-weight LLMs locally.
10 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running mlx-serve
| GPU | VRAM | Reports | Median t/s, Qwen3.8 125B · 6B active 3-bit |
|---|---|---|---|
| M4 Max 128GBapple | 128GB | 2 | no plain run of this model |
| M5 Max 128GBapple | 128GB | 2 | no plain run of this model |
| M5 Ultra 96GBapple | 96GB | 2 | no plain run of this model |
| M2 Max 96GBapple | 96GB | 1 | no plain run of this model |
| M3 Ultra 256GBapple | 256GB | 1 | no plain run of this model |
| M4 Max 64GBapple | 64GB | 1 | 52.6 |
| M5 Ultra 256GBapple | 256GB | 1 | no plain run of this model |
mlx-serve against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
No matched pair yet. No card has plain runs of one model size at one bit class on mlx-serve and on another engine, so llamaperf can't say how it compares on speed. Add a run.
mlx-serve results by GPU
Every card people have run mlx-serve on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
mlx-serve on M4 Max 128GB2 reports
- Qwen3.8 125B · 6B active · mixed 4-8bit60.0 t/s, 730 t/s prompt
- Qwen3.8 125B · 6B active · Q473.8 t/s(speculative)
mlx-serve on M5 Max 128GB2 reports
- Qwen3.8 125B · 6B active · Sushi-490.0 t/s, 1,900 t/s prompt(speculative)
- Qwen3.8 125B · 6B active · mixed-4-8bit40.0 t/s(speculative)
mlx-serve on M5 Ultra 96GB2 reports
- Qwen3.8 27B · 4.7bpw113.7 t/s, 3,191 t/s prompt(part in system RAM)
- Qwen3.8 125B · 6B active · 4-bit-8bit mix42.7 t/s, 828 t/s prompt(4 requests at once, speculative, part in system RAM)
mlx-serve on M2 Max 96GB1 report
- Qwen3.8 27B · 4bit20.1 t/s(speculative)
mlx-serve on M3 Ultra 256GB1 report
- Qwen3.8 27B · 4-bit84.1 t/s(speculative)
mlx-serve on M4 Max 64GB1 report
- Qwen3.8 125B · 6B active · iQ-MLX 3.3bpw52.6 t/s, 762 t/s prompt
mlx-serve on M5 Ultra 256GB1 report
- Qwen3.8 125B · 6B active · mixed 4/8-bit154.5 t/s, 4,183 t/s prompt(speculative)
Top models on mlx-serve
Frequently asked
Is mlx-serve faster than other engines?
llamaperf has no matched comparison for mlx-serve yet: no card has plain runs of the same model size at the same bit class on mlx-serve and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.