llamaperf
Oct 5, 2026
Throughput
117.0 t/s gen · 6066.0 t/s pp
KV cache
fp8
VRAM reported
64 GB

Summary

User reports DeepSeek-V4.1-Flash at 117 tok/s decode on one stream at 128k context and 6,066 tok/s prefill at 105k tokens on eight 64 GB CMP 170HX mining cards. Setup is vLLM with fp8 KV cache, PP=8, DSpark speculative decoding with 5 draft tokens, Engram tables in pinned host RAM, and 1M context. Decode drops to 96 tok/s at 512k. Eight concurrent streams reach 532 tok/s aggregate (66 per stream) at 105k. KV pool holds 6.17M tokens. One card was capped to 180 W after PCIe bus drops.