llamaperf
Oct 7, 2026
Throughput
60.0 t/s gen · 730.0 t/s pp
Quant
mixed 4-8bit (MLX)
System RAM
128 GB

Use cases

codingvisionlong-context

Summary

User reports Qwen3.8-Flash-Next at ~60 tok/s serial decode and ~730 tok/s prefill on an M4 Max 128GB. Setup is mlx-serve 26.8.11 with a mixed 4-bit/8-bit MLX pack, ~75 GB resident, prefix cache on, images included. With MTP enabled decode reaches 78 tok/s (+41% on code, a few percent slower on prose). A needle at 24.8k tokens was recovered with sparse attention engaged. The pack includes the MTP head and vision tower; the 51B n-gram table ships as a 32 GB mmapped file that is not resident. The author recommends a re-quantized iQ-MLX-4.7bpw version instead.