GLM-4.7 30B (3B active) Flash
on NVIDIA H100 80GB · SGLang
Oct 7, 2026
Summary
User reports GLM-4.7-Flash at 168 tok/s on a single NVIDIA H100 80GB, up from a 120 tok/s baseline with EAGLE3 speculative decoding.
Setup is SGLang v0.5.6 with the FlashInfer backend, a 277 MB EAGLE3 draft head, 6 draft tokens proposed per step, and a 40% acceptance rate averaging 2.4 accepted tokens per verification step.
At batch size 32 the same setup reaches 440 tok/s versus 259 tok/s baseline (1.70x), with per-request latency improving 1.30x. The user also reproduced the EAGLE3 paper's Llama-3.1-8B result at 1.25x versus the published 1.32x, and notes a SGLang token-counting bug that deflated measured throughput by about 35%.