GLM-5.3 320B (18B active) Flash
on NVIDIA RTX 3090 · Strata · 16,384 ctx
Oct 6, 2026
Summary
User reports GLM-5.3-Flash at 19.8 tok/s decode and 437.8 tok/s prefill on one RTX 3090 24 GB with 251 GB RAM.
Setup is the Strata engine with a UD-Q4_K_XL GGUF pack, --chunk 4096, 16K-token prompt, tiered expert cache streaming from RAM and NVMe.
Decode splits into 12.4 ms expert copies, 0.2 ms expert kernels and 37.2 ms dense/sync per 49.8 ms token. A 64K prompt prefills at 420.7 tok/s; a 1K prompt at 132 tok/s prefill and 19.8 tok/s decode; disk-only tier decodes at 1.15 tok/s. --chunk 8192 OOMs on 24 GB.