GLM-4.5-Air
on NVIDIA RTX Pro 6000 Blackwell · ik_llama.cpp · 16,384 ctx
Oct 5, 2026
Summary
User benchmarks GLM-4.5-Air with Unsloth UD-Q3_K_XL quant on an RTX PRO 6000 96GB with CPU offload for some experts, using ik_llama.cpp at 16k context with q8_0 KV cache.
Setup is ik_llama.cpp with UD-Q3_K_XL and q8_0 KV cache, offloading some expert layers to a Ryzen 9 7950X with 128GB DDR5-3600.
At 12288 tokens of context, prompt processing reaches 444.23 t/s and generation 8.25 t/s. The user also compares mainline llama.cpp and an Ubergarm IQ3_KT quant, noting a batch warm-up quirk in the forked llama-sweep-bench.