DeepSeek R1 671B (37B active)
on Unknown GPU · ik_llama.cpp
Oct 4, 2026
Summary
User benchmarks DeepSeek R1 671B at 9.9 t/s generation on a dual-socket Intel Xeon 6980P server with 1.5 TB RAM.
Setup is ik_llama.cpp with Q2_K_XL GGUF quant and F16 KV cache, running CPU-only on a single NUMA node with 43 threads.
The user compares llama.cpp (8.9 t/s) and ik_llama.cpp (9.9 t/s) for Q2_K_XL, and also tests Q4_K_M (10.0 t/s) and Q8_0 (7.5 t/s) with ik_llama.cpp. Dual-socket configurations degraded performance.