Qwen3 30B (3B active)
M5 Pro 24GB · llama.cpp
- reported speed:
- 73.9 tokens/s generation
- quant:
- IQ3_XXS (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports decode speed improvements on Qwen3-30B-A3B with IQ3_XXS quantization on an M5 MacBook Pro 24GB. The decode speed increased from 65.6 to 73.9 tok/s using a llama.cpp PR for Metal optimization. Setup uses llama.cpp with IQ3_XXS quantization. The user also tested on an M1 Pro 32GB. The user is looking for testers for Qwen3.8-Flash-Next and provides benchmark commands.