DeepSeek V2 16B (2.4B active) Lite
Unknown GPU · llama.cpp
- reported speed:
- 13.8 tokens/s generation
- quant:
- Q4_K_S
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek-V2-Lite-Chat Q4_K_S at 13.79 tok/s with llama.cpp on an Intel Core i5-11300H CPU at 4 threads. The user's own C99 engine runs the same model at 1.90 tok/s on the same hardware, and the post is about closing that gap.