llamaperf
Sep 30, 2026
Throughput
63.7 t/s gen
Quant
NVFP4 (NVFP4)
KV cache
FP8

Summary

User reports GLM-5.3-Flash at 63.7 tok/s single-request decode on 4x RTX PRO 6000 Blackwell GPUs. Setup is an experimental pinned SGLang build with an NVFP4 checkpoint, FP8 E4M3 KV cache, FlashInfer sparse MLA, and 262,144-token context, using tensor parallel TP4/EP4. With EAGLE speculative decoding (adaptive, up to 5 steps) the single-request figure rises to 82.8 tok/s. Eight concurrent requests give 183.8 tok/s aggregate baseline and 182.1 tok/s with EAGLE.