DeepSeek V2 16B (2.4B active) Lite
on Unknown GPU · llama.cpp
Oct 8, 2026
Summary
User reports DeepSeek-V2-Lite-Chat Q4_K_S at 13.79 tok/s with llama.cpp on an Intel Core i5-11300H CPU at 4 threads.
The user's own C99 engine runs the same model at 1.90 tok/s on the same hardware, and the post is about closing that gap.