GLM-5.3 320B (18B active) Flash
on M5 Ultra 512GB · mlx-vlm · 131,072 ctx
Sep 25, 2026
Summary
User reports GLM-flash-4bit with MTP on an M5 Ultra 512GB, reaching 735.5 prompt t/s and 71.1 generation t/s at 131,072 prompt tokens.
Setup is raw mlx-vlm with --prefill-step-size 8192. At 32,768 tokens the run gave 1056.0 prompt t/s and 72.7 generation t/s; at 65,536 tokens it gave 920.0 prompt t/s and 73.7 generation t/s.
With --prefill-step-size 2048 the same trials gave 623.6 prompt t/s and 50.8 generation t/s at 131,072 tokens. The user also notes omlx with adaptive step size matched 8k performance from 2048 at 64k context and above.