llamaperf

Qwen3 30B (3B active)

on M5 Pro 24GB · llama.cpp

Tone: positive
Oct 6, 2026
Throughput
73.9 t/s gen
Quant
IQ3_XXS (GGUF)
System RAM
24 GB

Summary

User reports decode speed improvements on Qwen3-30B-A3B with IQ3_XXS quantization on an M5 MacBook Pro 24GB. The decode speed increased from 65.6 to 73.9 tok/s using a llama.cpp PR for Metal optimization. Setup uses llama.cpp with IQ3_XXS quantization. The user also tested on an M1 Pro 32GB. The user is looking for testers for Qwen3.8-Flash-Next and provides benchmark commands.