llamaperf
Sep 24, 2026
Throughput
95.0 t/s gen
Quant
FP8
KV cache
FP8
VRAM reported
256 GB

Use cases

codinglong-context

Summary

User reports DeepSeek V4 Flash Vision at up to ~95 tok/s on 4× CMP 170HX 64GB (256GB aggregate HBM2e). Setup is vLLM with pipeline parallel 4, max_model_len 262144, FP8 KV cache, DSpark speculative decoding with 6 tokens, and max_num_seqs 2. Normal coding workload runs ~50–70 tok/s, strong speculative-decoding periods ~80–90 tok/s, peak ~95 tok/s. Cards sit around 56–62GB VRAM each at 50–60°C. Next tests planned for Qwen3.8 Flash Next and GLM 5.3.