llamaperf
Oct 7, 2026
Throughput
168.0 t/s gen
VRAM reported
80 GB

Summary

User reports GLM-4.7-Flash at 168 tok/s on a single NVIDIA H100 80GB, up from a 120 tok/s baseline with EAGLE3 speculative decoding. Setup is SGLang v0.5.6 with the FlashInfer backend, a 277 MB EAGLE3 draft head, 6 draft tokens proposed per step, and a 40% acceptance rate averaging 2.4 accepted tokens per verification step. At batch size 32 the same setup reaches 440 tok/s versus 259 tok/s baseline (1.70x), with per-request latency improving 1.30x. The user also reproduced the EAGLE3 paper's Llama-3.1-8B result at 1.25x versus the published 1.32x, and notes a SGLang token-counting bug that deflated measured throughput by about 35%.