llamaperf

DeepSeek R1

DeepSeek · 1 report

Thin page (1 of 3 reports needed for indexing). Add yours.
reported speed:
2.0 tokens/s generation
quant:
Q3 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports about 2 t/s running DeepSeek-R1-0528 Q3 GGUF on 2x RTX 3090 with 48 GB VRAM and 512 GB DDR5 system RAM. No engine is named. The user asks for upgrade advice to reach 15+ t/s with a very large model, and asks about Threadripper, EPYC and Xeon platforms.

Sep 11, 2026