llamaperf

DeepSeek V4 Pro

DeepSeek · 2 reports

Thin page (2 of 3 reports needed for indexing). Add yours.

DeepSeek V4 Pro

RTX PRO 6000 Max-Q · llama.cpp · 1,048,576 ctx

Benchmark of DeepSeek V4 Pro GGUF (794GB) on llama.cpp branch with expert offloading. Hardware: Epyc 9374F, 12x96GB DDR5, RTX PRO 6000 Max-Q. Prompt processing speeds range from 192 t/s (8K context) to 66 t/s (1M context). Generation speeds range from 11.73 t/s to 5.83 t/s. RAM usage 69.3% of 1152GB, VRAM usage 78986MiB of 96GB. Power ~500W during PP. Notes on mainline llama.cpp issues: memory waste, broken quantized KV cache, bugs with prompt cache reuse.

DeepSeek V4 Pro

RTX PRO 6000 Max-Q · ktransformers · 65,536 ctx

reported speed:
7.1 tokens/s generation · 46.2 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark at various context depths. GPU VRAM usage 90815MiB/97887MiB. RAM usage 907.5GB/1152GB. CPU: Epyc 9374F. Power: GPU ~100W PP, ~150W TG; CPU+MB ~400W. Original model files, no conversion.