llamaperf

GLM-5.3 320B (18B active) Flash

on 2× NVIDIA CMP 170HX 64GB (unlocked) · ExLlamaV3 · 8,192 ctx

Tone: positive
Oct 8, 2026
Throughput
95.6 t/s gen · 1529.0 t/s pp
Quant
3.05bpw (EXL3)
KV cache
Q8
System RAM
80 GB
VRAM reported
128 GB

Use cases

codingagenticlong-context

Summary

User reports GLM-5.3-Flash at 95.6 t/s decode and 1,529 t/s prefill on an 8K input, running on two CMP 170HX 64GB cards. Setup is ExLlamaV3 1.5.4 with an EXL3 3.05bpw target, DFlash2 EXL3 6bpw K7 drafter and Q8 KV cache, target weights fully resident in HBM at about 116.6GiB across the two cards. Decode stays near 90 t/s out to a 385K input (90.1 t/s, 1,535 t/s prefill); DFlash acceptance was around 87%, against roughly 60 t/s on the older MTP d2 path. The user also compared a Qwen3.8-Flash-Next AWQ INT4 + FP8 PLE vLLM setup on the same machine, which reached 129.5 t/s decode and 5,513 t/s prefill at 8K, but GLM finished three one-shot coding tasks sooner (4m40s vs 5m30s, 10m05s vs 14m40s, 4m40s vs 24m10s).