Qwen3 30B (3B active)
on M4 Max 128GB · vllm-mlx
Sep 24, 2026
Summary
User reports Qwen3-30B-A3B 4-bit at 127.7 tok/s on an M4 Max 128GB.
Setup is vllm-mlx, a native MLX inference server, with greedy decoding and single-stream decode.
Two other models were also measured on the same machine: Qwen3-0.6B 8-bit at 417.9 tok/s and Llama-3.2-3B-Instruct 4-bit at 205.6 tok/s. Continuous-batching results, KV-cache quantisation and MoE top-k sweeps are documented separately.