GLM-4 9B
on Intel Arc B580 12GB · llama.cpp · 2,048 ctx
Oct 6, 2026
Summary
User reports GLM-4 9B at 21.25 t/s on an Intel Arc B580 12GB, with GPU utilization near 40%.
Setup is llama.cpp (ipex-llm[cpp] build 2024.12.17) with Q4_0 GGUF and F16 KV cache, 2048 context, all 41 layers offloaded to the GPU.
The user considers the performance low for the hardware. The prompt eval figure of 41.47 t/s is over only 5 tokens and is not reported as a prefill rate.