llamaperf

DeepSeek V2 16B (2.4B active) Lite

on Unknown GPU · llama.cpp

Tone: mixed
Oct 8, 2026
Throughput
13.8 t/s gen
Quant
Q4_K_S

Summary

User reports DeepSeek-V2-Lite-Chat Q4_K_S at 13.79 tok/s with llama.cpp on an Intel Core i5-11300H CPU at 4 threads. The user's own C99 engine runs the same model at 1.90 tok/s on the same hardware, and the post is about closing that gap.