llamaperf

Phi-3.5 3.8B mini

on M4 Pro 64GB · mlx-lm · 4,096 ctx

Tone: positive
Sep 26, 2026
Throughput
72.2 t/s gen
Quant
4bit (MLX)
KV cache
kv4
System RAM
64 GB

Summary

User reports Phi-3.5-mini-instruct-4bit at 72.2 t/s on a Mac Mini M4 Pro 64GB. Setup is mlx-lm 0.30.7 with 4-bit weights and kv4 KV-cache quantization at 4096 generated tokens. KV-cache quantization with kv4 uses 0.51 GB and is 1.1% faster than unquantized (71.4 t/s), while kv8 uses 0.91 GB and is 4.9% slower (67.9 t/s). User also tested Qwen2.5-14B-Instruct-4bit and found GQA halves KV-cache per token versus Phi-3.5-mini's MHA.