Qwen3.8 125B (6B active) Flash-Next
on M5 Max 128GB · mlx-serve
Sep 29, 2026
Summary
User reports Qwen3.8-Flash-Next at 1900 t/s prefill and 90 t/s generation on an M5 Max 128GB.
Setup is mlx-serve with Sushi-4 quant, an MLX-EXL3 hybrid, with MTP and Vision supported.
User also cites M1 Max 64GB at ~350 t/s prefill and 30 t/s gen, and M5 Pro at ~900 t/s prefill and 60 t/s gen.