GLM-5.3 320B (18B active) Flash
on 2× NVIDIA CMP 170HX 64GB (unlocked) · ExLlamaV3 · 8,192 ctx
Oct 8, 2026
Use cases
codingagenticlong-context
Summary
User reports GLM-5.3-Flash at 95.6 t/s decode and 1,529 t/s prefill on an 8K input, running on two CMP 170HX 64GB cards.
Setup is ExLlamaV3 1.5.4 with an EXL3 3.05bpw target, DFlash2 EXL3 6bpw K7 drafter and Q8 KV cache, target weights fully resident in HBM at about 116.6GiB across the two cards.
Decode stays near 90 t/s out to a 385K input (90.1 t/s, 1,535 t/s prefill); DFlash acceptance was around 87%, against roughly 60 t/s on the older MTP d2 path. The user also compared a Qwen3.8-Flash-Next AWQ INT4 + FP8 PLE vLLM setup on the same machine, which reached 129.5 t/s decode and 5,513 t/s prefill at 8K, but GLM finished three one-shot coding tasks sooner (4m40s vs 5m30s, 10m05s vs 14m40s, 4m40s vs 24m10s).