DeepSeek V4 Pro
RTX PRO 6000 Max-Q · llama.cpp · 1,048,576 ctx
Benchmark of DeepSeek V4 Pro GGUF (794GB) on llama.cpp branch with expert offloading. Hardware: Epyc 9374F, 12x96GB DDR5, RTX PRO 6000 Max-Q. Prompt processing speeds range from 192 t/s (8K context) to 66 t/s (1M context). Generation speeds range from 11.73 t/s to 5.83 t/s. RAM usage 69.3% of 1152GB, VRAM usage 78986MiB of 96GB. Power ~500W during PP. Notes on mainline llama.cpp issues: memory waste, broken quantized KV cache, bugs with prompt cache reuse.