llamaperf
Sep 24, 2026
Throughput
127.7 t/s gen
Quant
4-bit (MLX)
System RAM
128 GB

Summary

User reports Qwen3-30B-A3B 4-bit at 127.7 tok/s on an M4 Max 128GB. Setup is vllm-mlx, a native MLX inference server, with greedy decoding and single-stream decode. Two other models were also measured on the same machine: Qwen3-0.6B 8-bit at 417.9 tok/s and Llama-3.2-3B-Instruct 4-bit at 205.6 tok/s. Continuous-batching results, KV-cache quantisation and MoE top-k sweeps are documented separately.