llamaperf
Oct 3, 2026
Throughput
125.8 t/s gen · 363.9 t/s pp
Quant
4bit (MLX)
System RAM
24 GB

Use cases

codingagentic

Summary

User benchmarks DeepSeek-Coder-V2-Lite-Instruct at 125.8 tok/s generation and 363.9 tok/s prompt processing on an Apple M5 Pro with 24GB unified memory. Setup is MLX with a 4bit quant (8.2GB) on a single machine, running a fixed coding task capped at 1500 output tokens. The model produced correct working code; the user notes that 24GB has a hard ceiling around 20GB models and that background downloads measurably affect tok/s.