llamaperf

GLM-5.3

Zhipu AI · 5 reports

reported speed:
1.8 tokens/s generation
quant:
JANG

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Benchmark of MoE streaming PR for oMLX on M4 Pro 48GB. Four models tested: GLM-5.3-Flash-JANG-MTP (10.52 GiB after load, 14.68 GiB peak, 12.36s TTFT, 1.82 tok/s), Qwen3.8-JANG 4S (7.04 GiB, 11.19 GiB, 8.56s, 3.63 tok/s), Qwen3.8-JANG 4M (7.05 GiB, 11.11 GiB, 11.23s, 3.19 tok/s), DeepSeek-V4-Flash-0731-JANG (8.25 GiB, 16.87 GiB, 6.52s, 2.71 tok/s). Primary record uses GLM-5.3-Flash-JANG-MTP. Other models: Qwen3.8 (JANG 4S and 4M quants), DeepSeek V4 Flash (0731 variant, JANG quant). MoE streaming allows running larger MoE models with lower memory footprint at the cost of speed.

Tone: positive
reported speed:
37.4 tokens/s generation · 550.0 tokens/s prompt processing
quant:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcodingtool-use

Optimized GLM-5.3-Flash on M3 Ultra. Headline 60tps for SQL, 38tps average with agent harness. Prefill 550 t/s at 62k, generation 37.4 t/s at 300k context. Uses dflash drafter for speculative decoding. Accuracy preserved.

GLM-5.3 Flash

RTX 5090 · LayerStoRm · 1,000,000 ctx

Tone: positive
reported speed:
24.5 tokens/s generation · 159.0 tokens/s prompt processing
quant:
UD-Q4_K_XL

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Runs 186 GiB MoE on 96 GB VRAM using RAM for pinned experts. Decode 24.5 tok/s @8k, 27.0 tok/s @0k. Prefill 159 tok/s @27k. Prefix caching with mid-prompt checkpoints reduces TTFT from 67.5s to 18.4s at 8k, and ~923s to 79s at 97k. Host RAM ~208 GB pinned. CPU does no compute, only feeds experts. NUMA-aware transfers. NVIDIA SM120 only.

reported speed:
55.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Post discusses GLM-5.3 pricing and intelligence, but does not run it. Mentions running Qwen3.8-27B on RTX 3090 with IQ4_XS quant, generation speed 55 tok/s. Also mentions DeepSeek V4 Flash and Strix Halo hardware.

Tone: positive

User reports GLM 5.3 Flash support in ds4 branch, running on M4 Max 128GB. No performance numbers provided.