llamaperf
Sep 29, 2026
Throughput
90.0 t/s gen · 1900.0 t/s pp
Quant
Sushi-4 (MLX)
System RAM
128 GB

Summary

User reports Qwen3.8-Flash-Next at 1900 t/s prefill and 90 t/s generation on an M5 Max 128GB. Setup is mlx-serve with Sushi-4 quant, an MLX-EXL3 hybrid, with MTP and Vision supported. User also cites M1 Max 64GB at ~350 t/s prefill and 30 t/s gen, and M5 Pro at ~900 t/s prefill and 60 t/s gen.