Phi-3.5 3.8B mini
M4 Pro 64GB · mlx-lm · 4,096 ctx
- reported speed:
- 72.2 tokens/s generation
- quant:
- 4bit (MLX)
- kv:
- kv4
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Phi-3.5-mini-instruct-4bit at 72.2 t/s on a Mac Mini M4 Pro 64GB. Setup is mlx-lm 0.30.7 with 4-bit weights and kv4 KV-cache quantization at 4096 generated tokens. KV-cache quantization with kv4 uses 0.51 GB and is 1.1% faster than unquantized (71.4 t/s), while kv8 uses 0.91 GB and is 4.9% slower (67.9 t/s). User also tested Qwen2.5-14B-Instruct-4bit and found GQA halves KV-cache per token versus Phi-3.5-mini's MHA.