llamaperf
Sep 23, 2026
Throughput
98.1 t/s gen · 5300.0 t/s pp
Quant
MXFP4+FP8
VRAM reported
64 GB

Summary

User reports DeepSeek-V4-Flash-0731 at 98.1 t/s single-stream decode on 4x CMP 170HX (64 GB HBM2e each), up from a 50.8 t/s baseline with DSpark speculative decoding. Setup is vLLM with pipeline parallelism, native MXFP4+FP8 weights (~155.4 GiB), context verified to 1,047,736 tokens. At 100k context single-stream decode is 38.8 t/s with 14.6 s time to first token. Aggregate throughput across 64 concurrent requests is 712.8 t/s with DSpark (472.0 t/s baseline), and 90.0 t/s at 100k context. Prefill across the 25k-77k context range is about 5,300 t/s.