llamaperf

DeepSeek V2

1 report

DeepSeek V2 VRAM requirements by size and quant →

How does DeepSeek V2 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for DeepSeek V2 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run DeepSeek V2 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for DeepSeek V2

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

DeepSeek V2 16B (2.4B active) Lite

Unknown GPU · llama.cpp

Tone: mixed
reported speed:
13.8 tokens/s generation
quant:
Q4_K_S

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V2-Lite-Chat Q4_K_S at 13.79 tok/s with llama.cpp on an Intel Core i5-11300H CPU at 4 threads. The user's own C99 engine runs the same model at 1.90 tok/s on the same hardware, and the post is about closing that gap.

Oct 8, 2026

Get a weekly email of new DeepSeek V2 reports on any GPU.

Email me new reports