llamaperf

DeepSeek-Coder-V2-Lite

1 report

As of 7 Oct 2026, DeepSeek-Coder-V2-Lite 16B · 2.4B active at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

DeepSeek-Coder-V2-Lite VRAM requirements by size and quant →

How does DeepSeek-Coder-V2-Lite run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for DeepSeek-Coder-V2-Lite on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run DeepSeek-Coder-V2-Lite yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for DeepSeek-Coder-V2-Lite

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: positive
reported speed:
125.8 tokens/s generation · 363.9 tokens/s prompt processing
quant:
4bit (MLX)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User benchmarks DeepSeek-Coder-V2-Lite-Instruct at 125.8 tok/s generation and 363.9 tok/s prompt processing on an Apple M5 Pro with 24GB unified memory. Setup is MLX with a 4bit quant (8.2GB) on a single machine, running a fixed coding task capped at 1500 output tokens. The model produced correct working code; the user notes that 24GB has a hard ceiling around 20GB models and that background downloads measurably affect tok/s.

Oct 3, 2026

Get a weekly email of new DeepSeek-Coder-V2-Lite reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
DeepSeek-Coder-V2-Lite 16B (2.4B active)
M5 Pro 24GB
4bit
MLX
Not reported125.8 tokens/s