GLM-5.3 320B (18B active) Flash
on 4× NVIDIA RTX Pro 6000 Blackwell · SGLang · 262,144 ctx
Sep 30, 2026
Summary
User reports GLM-5.3-Flash at 63.7 tok/s single-request decode on 4x RTX PRO 6000 Blackwell GPUs.
Setup is an experimental pinned SGLang build with an NVFP4 checkpoint, FP8 E4M3 KV cache, FlashInfer sparse MLA, and 262,144-token context, using tensor parallel TP4/EP4.
With EAGLE speculative decoding (adaptive, up to 5 steps) the single-request figure rises to 82.8 tok/s. Eight concurrent requests give 183.8 tok/s aggregate baseline and 182.1 tok/s with EAGLE.