GLM-4 9B
Intel Arc B580 12GB · llama.cpp · 2,048 ctx
- reported speed:
- 21.3 tokens/s generation
- quant:
- Q4_0 (GGUF)
- kv:
- F16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports GLM-4 9B at 21.25 t/s on an Intel Arc B580 12GB, with GPU utilization near 40%. Setup is llama.cpp (ipex-llm[cpp] build 2024.12.17) with Q4_0 GGUF and F16 KV cache, 2048 context, all 41 layers offloaded to the GPU. The user considers the performance low for the hardware. The prompt eval figure of 41.47 t/s is over only 5 tokens and is not reported as a prefill rate.