GLM-5.3 320B (18B active) Flash
M5 Ultra 256GB · oMLX · 4,096 ctx
- reported speed:
- 76.7 tokens/s generation · 2270.0 tokens/s prompt processing
- quant:
- oQ4 (MLX)
- kv:
- 4-bit
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks oMLX 0.7.0 against 0.7.0rc1 on an M5 Ultra 256GB, running GLM-5.3-Flash oQ4 at 4K–200K context. Prefill improves from 792 to 2,270 tok/s (median) and decode from 55.4 to 76.7 tok/s. A 1M-token prompt completes with prefill 1,485 tok/s, decode 41.5 tok/s, and time to first token 11.8 min. Setup uses Lightning MTP, TurboQuant KV 4-bit, and one model loaded at a time.