llamaperf
Oct 6, 2026
Throughput
200.9 t/s gen · 1943.0 t/s pp
Quant
FP4 (FP8)
System RAM
128 GB
VRAM reported
96 GB

Use cases

long-contexttool-usevision

Summary

User reports DeepSeek-V4.1-Flash at 200.9 t/s decode and 1,943.0 t/s prefill on 4× RTX PRO 6000 Blackwell 96GB GPUs. Setup is SGLang with FP4 routed experts and FP8 components, DSpark block 5 speculative decoding, 524,288 context, 64 GiB DDR5 cache, NVMe offload, 8 slots, GPUs capped at 275 W over PCIe without NVLink. The 200.9 t/s is the C1 single-stream figure from a 45-case sweep; total decode across 8 concurrent streams reached 713.5 t/s. The 4M populated-KV gate passed with 4,063,744 tokens across eight 500k inputs. Long-context answer quality and a one-hour soak remain unqualified.