llamaperf
Oct 6, 2026
Throughput
19.8 t/s gen · 437.8 t/s pp
Quant
UD-Q4_K_XL (GGUF)
System RAM
251 GB
VRAM reported
24 GB

Summary

User reports GLM-5.3-Flash at 19.8 tok/s decode and 437.8 tok/s prefill on one RTX 3090 24 GB with 251 GB RAM. Setup is the Strata engine with a UD-Q4_K_XL GGUF pack, --chunk 4096, 16K-token prompt, tiered expert cache streaming from RAM and NVMe. Decode splits into 12.4 ms expert copies, 0.2 ms expert kernels and 37.2 ms dense/sync per 49.8 ms token. A 64K prompt prefills at 420.7 tok/s; a 1K prompt at 132 tok/s prefill and 19.8 tok/s decode; disk-only tier decodes at 1.15 tok/s. --chunk 8192 OOMs on 24 GB.