llamaperf

oMLX

An inference engine for running open-weight LLMs locally.

23 community reports

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running oMLX

GPUVRAMReportsMedian t/s, Qwen3.8 27B 4-bit
M5 Ultra 256GBapple256GB750.0
M4 Max 128GBapple128GB6no plain run of this model
M3 Ultra 512GBapple512GB4no plain run of this model
M4 Pro 48GBapple48GB2no plain run of this model
M1 Ultra 64GBapple64GB1no plain run of this model
M2 Max 96GBapple96GB1no plain run of this model
M3 Max 128GBapple128GB1no plain run of this model
M3 Max 48GBapple48GB1no plain run of this model

oMLX against other engines

Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.

No matched pair yet. No card has plain runs of one model size at one bit class on oMLX and on another engine, so llamaperf can't say how it compares on speed. Add a run.

oMLX results by GPU

Every card people have run oMLX on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.

oMLX on M4 Max 128GB4 reports

oMLX on M4 Pro 48GB2 reports

oMLX on M2 Max 96GB1 report

oMLX on M3 Max 48GB1 report

oMLX on M3 Max 96GB1 report

oMLX on M3 Ultra 256GB1 report

oMLX on M5 Max 64GB1 report

oMLX on M5 Pro 64GB1 report

oMLX, GPU not identified2 reports

Top models on oMLX

Frequently asked

Is oMLX faster than other engines?

llamaperf has no matched comparison for oMLX yet: no card has plain runs of the same model size at the same bit class on oMLX and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.