llamaperf

GLM-4.5-Air

on NVIDIA RTX Pro 6000 Blackwell · ik_llama.cpp · 16,384 ctx

Oct 5, 2026
Throughput
8.3 t/s gen · 444.2 t/s pp
Quant
UD-Q3_K_XL (GGUF)
KV cache
q8_0
System RAM
128 GB
VRAM reported
96 GB

Summary

User benchmarks GLM-4.5-Air with Unsloth UD-Q3_K_XL quant on an RTX PRO 6000 96GB with CPU offload for some experts, using ik_llama.cpp at 16k context with q8_0 KV cache. Setup is ik_llama.cpp with UD-Q3_K_XL and q8_0 KV cache, offloading some expert layers to a Ryzen 9 7950X with 128GB DDR5-3600. At 12288 tokens of context, prompt processing reaches 444.23 t/s and generation 8.25 t/s. The user also compares mainline llama.cpp and an Ubergarm IQ3_KT quant, noting a batch warm-up quirk in the forked llama-sweep-bench.