llamaperf
Oct 6, 2026
Throughput
76.7 t/s gen · 2270.0 t/s pp
Quant
oQ4 (MLX)
KV cache
4-bit
System RAM
256 GB

Summary

User benchmarks oMLX 0.7.0 against 0.7.0rc1 on an M5 Ultra 256GB, running GLM-5.3-Flash oQ4 at 4K–200K context. Prefill improves from 792 to 2,270 tok/s (median) and decode from 55.4 to 76.7 tok/s. A 1M-token prompt completes with prefill 1,485 tok/s, decode 41.5 tok/s, and time to first token 11.8 min. Setup uses Lightning MTP, TurboQuant KV 4-bit, and one model loaded at a time.