- reported speed:
- 2.0 tokens/s generation
- quant:
- Q3 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports about 2 t/s running DeepSeek-R1-0528 Q3 GGUF on 2x RTX 3090 with 48 GB VRAM and 512 GB DDR5 system RAM. No engine is named. The user asks for upgrade advice to reach 15+ t/s with a very large model, and asks about Threadripper, EPYC and Xeon platforms.