- reported speed:
- 125.8 tokens/s generation · 363.9 tokens/s prompt processing
- quant:
- 4bit (MLX)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User benchmarks DeepSeek-Coder-V2-Lite-Instruct at 125.8 tok/s generation and 363.9 tok/s prompt processing on an Apple M5 Pro with 24GB unified memory. Setup is MLX with a 4bit quant (8.2GB) on a single machine, running a fixed coding task capped at 1500 output tokens. The model produced correct working code; the user notes that 24GB has a hard ceiling around 20GB models and that background downloads measurably affect tok/s.