Qwen3 30B (3B active)
on M5 Pro 24GB · llama.cpp
Oct 6, 2026
Summary
User reports decode speed improvements on Qwen3-30B-A3B with IQ3_XXS quantization on an M5 MacBook Pro 24GB. The decode speed increased from 65.6 to 73.9 tok/s using a llama.cpp PR for Metal optimization.
Setup uses llama.cpp with IQ3_XXS quantization. The user also tested on an M1 Pro 32GB.
The user is looking for testers for Qwen3.8-Flash-Next and provides benchmark commands.