DeepSeek V4.1 Flash 552B (16B active)
8× NVIDIA CMP 170HX 64GB (unlocked) · vLLM · 1,048,576 ctx
- reported speed:
- 117.0 tokens/s generation · 6066.0 tokens/s prompt processing
- kv:
- fp8
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek-V4.1-Flash at 117 tok/s decode on one stream at 128k context and 6,066 tok/s prefill at 105k tokens on eight 64 GB CMP 170HX mining cards. Setup is vLLM with fp8 KV cache, PP=8, DSpark speculative decoding with 5 draft tokens, Engram tables in pinned host RAM, and 1M context. Decode drops to 96 tok/s at 512k. Eight concurrent streams reach 532 tok/s aggregate (66 per stream) at 105k. KV pool holds 6.17M tokens. One card was capped to 180 W after PCIe bus drops.