llamaperf

GLM-4

1 report

As of 7 Oct 2026, GLM-4 9B at 4-bit on the hardware it is most run on, with the median of plain runs (one device, one request, no speculative decoding, the whole model in its memory):

GLM-4 VRAM requirements by size and quant →

How does GLM-4 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for GLM-4 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run GLM-4 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for GLM-4

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

GLM-4 9B

Intel Arc B580 12GB · llama.cpp · 2,048 ctx

Tone: negative
reported speed:
21.3 tokens/s generation
quant:
Q4_0 (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports GLM-4 9B at 21.25 t/s on an Intel Arc B580 12GB, with GPU utilization near 40%. Setup is llama.cpp (ipex-llm[cpp] build 2024.12.17) with Q4_0 GGUF and F16 KV cache, 2048 context, all 41 layers offloaded to the GPU. The user considers the performance low for the hardware. The prompt eval figure of 41.47 t/s is over only 5 tokens and is not reported as a prefill rate.

Oct 6, 2026

Get a weekly email of new GLM-4 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
GLM-4 9B
Intel Arc B580 12GB
Q4_0
llama.cpp
2,04821.3 tokens/s