- quant:
- Q3 (gguf)
User reports running a frontier model on a home PC with 24 GB of VRAM. The user notes the run is slow.
DeepSeek · 3 reports
User reports running a frontier model on a home PC with 24 GB of VRAM. The user notes the run is slow.
User is setting up a 16x DGX Spark cluster to run frontier models locally. User mentions DeepSeek V4 pro, Kimi K3, GLM 5.5, and Minimax M4 as future models. No benchmark numbers are provided.
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports decode speeds at various context depths: 28 t/s at the start, 23.5 t/s at 45k, and 18 t/s at 192k. The run was maintained with 8k token output. Prefill performance is mentioned but no numbers are given.