Gemma 4M5 32GB · MLX · 130,173 ctxthroughput: 3029.0 t/s ppUser built custom w8a8 kernels for M5 MacBook Air. Baseline prefill 2193 tps, improved to 3029 tps for 130k tokens. Also mentions llama.cpp for Macs.