ik_llama.cpp
An inference engine for running open-weight LLMs locally.
11 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running ik_llama.cpp
| GPU | VRAM | Reports | Median t/s, Qwen3.8 125B · 6B active 4-bit |
|---|---|---|---|
| NVIDIA RTX 3090nvidia | 24GB | 3 | no plain run of this model |
| NVIDIA RTX Pro 6000 Blackwellnvidia | 96GB | 2 | 40.0 |
| NVIDIA RTX 3060 12GBnvidia | 12GB | 1 | no plain run of this model |
| NVIDIA RTX 4070nvidia | 12GB | 1 | no plain run of this model |
| NVIDIA RTX 4070 Supernvidia | 12GB | 1 | no plain run of this model |
| NVIDIA RTX 5090nvidia | 32GB | 1 | no plain run of this model |
| NVIDIA T4 16GBnvidia | 16GB | 1 | no plain run of this model |
| NVIDIA V100 16GBnvidia | 16GB | 1 | no plain run of this model |
ik_llama.cpp against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
| Model and GPU | ik_llama.cpp | Other engine |
|---|---|---|
| Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell | 40.0 t/sone run, 200K context | llama.cpp105.5 t/smedian of 2, 2K to 240K context |
| Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell | 40.0 t/sone run, 200K context | NInfer172.0 t/sone run, 512 context |
ik_llama.cpp results by GPU
Every card people have run ik_llama.cpp on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
ik_llama.cpp on NVIDIA RTX 30902 reports
- Qwen3.6 27B · IQ4_KS72.0 t/s(speculative)
- Qwen3.6 27B · IQ4_KS72.9 t/s, 1,261 t/s prompt(speculative)
ik_llama.cpp on NVIDIA RTX Pro 6000 Blackwell2 reports
- GLM-4.5-Air · UD-Q3_K_XL8.3 t/s, 444 t/s prompt(part in system RAM)
- Qwen3.8 125B · 6B active · IQ440.0 t/s, 2,000 t/s prompt
ik_llama.cpp on NVIDIA RTX 3060 12GB1 report
- Qwen3.8 125B · 6B active · Q4_K_XL22.1 t/s, 66 t/s prompt(speculative, part in system RAM)
ik_llama.cpp on NVIDIA RTX 40701 report
- Qwen3.6 35B · 3B active · Q3_K_P28.7 t/s(part in system RAM)
ik_llama.cpp on NVIDIA RTX 4070 Super1 report
- Qwen3.6 35B · 3B active · IQ4_XS-4.19bpw110.2 t/s(speculative, part in system RAM)
ik_llama.cpp on NVIDIA RTX 50901 report
- Qwen3.8 125B · 6B active · AP Q4_K_M38.8 t/s, 200 t/s prompt(part in system RAM)
ik_llama.cpp on NVIDIA T4 16GB1 report
- Qwen3.8 125B · 6B active · UD-Q4_K_XL17.6 t/s, 160 t/s prompt(part in system RAM)
ik_llama.cpp on NVIDIA V100 16GB1 report
- Qwen3.5 9B · Q6_K4.6 t/s, 53 t/s prompt
ik_llama.cpp, GPU not identified1 report
Top models on ik_llama.cpp
Frequently asked
Is ik_llama.cpp faster than other engines?
These counts come from reports with the same card, model size and bit class on both engines, plain runs only. Against llama.cpp there is one matched comparison, and llama.cpp is faster. Against NInfer there is one matched comparison, and NInfer is faster. Context length and engine build still differ between the runs, so each comparison is a single data point.