DeepSeek V4.1 Flash 552B (16B active)
4× NVIDIA RTX Pro 6000 Blackwell · SGLang · 524,288 ctx
- reported speed:
- 200.9 tokens/s generation · 1943.0 tokens/s prompt processing
- quant:
- FP4 (FP8)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports DeepSeek-V4.1-Flash at 200.9 t/s decode and 1,943.0 t/s prefill on 4× RTX PRO 6000 Blackwell 96GB GPUs. Setup is SGLang with FP4 routed experts and FP8 components, DSpark block 5 speculative decoding, 524,288 context, 64 GiB DDR5 cache, NVMe offload, 8 slots, GPUs capped at 275 W over PCIe without NVLink. The 200.9 t/s is the C1 single-stream figure from a 45-case sweep; total decode across 8 concurrent streams reached 713.5 t/s. The 4M populated-KV gate passed with 4,063,744 tokens across eight 500k inputs. Long-context answer quality and a one-hour soak remain unqualified.