llamaperf
Sep 27, 2026
Throughput
52.6 t/s gen · 762.0 t/s pp
Quant
iQ-MLX 3.3bpw (MLX)
System RAM
64 GB

Use cases

codingagenticvisionlong-context

Summary

User reports Qwen3.8-Flash-Next at 52.6 tok/s decode and 762 tok/s prefill on an M4 Max 64GB. Setup is mlx-serve with an iQ-MLX 3.3bpw pack, 64k context, 52 GB resident; the 32 GB n-gram table is mmapped and not resident. The same run measured the mixed-4-8bit pack at 55.5 tok/s decode, 754 tok/s prefill and 278 ms TTFT; MTP is in the pack but default-off and not re-measured.