Phi-3.5 3.8B mini
on M4 Pro 64GB · mlx-lm · 4,096 ctx
Sep 26, 2026
Summary
User reports Phi-3.5-mini-instruct-4bit at 72.2 t/s on a Mac Mini M4 Pro 64GB.
Setup is mlx-lm 0.30.7 with 4-bit weights and kv4 KV-cache quantization at 4096 generated tokens.
KV-cache quantization with kv4 uses 0.51 GB and is 1.1% faster than unquantized (71.4 t/s), while kv8 uses 0.91 GB and is 4.9% slower (67.9 t/s). User also tested Qwen2.5-14B-Instruct-4bit and found GQA halves KV-cache per token versus Phi-3.5-mini's MHA.