llamaperf

ik_llama.cpp

An inference engine for running open-weight LLMs locally.

11 community reports

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running ik_llama.cpp

GPUVRAMReportsMedian t/s, Qwen3.8 125B · 6B active 4-bit
NVIDIA RTX 3090nvidia24GB3no plain run of this model
NVIDIA RTX Pro 6000 Blackwellnvidia96GB240.0
NVIDIA RTX 3060 12GBnvidia12GB1no plain run of this model
NVIDIA RTX 4070nvidia12GB1no plain run of this model
NVIDIA RTX 4070 Supernvidia12GB1no plain run of this model
NVIDIA RTX 5090nvidia32GB1no plain run of this model
NVIDIA T4 16GBnvidia16GB1no plain run of this model
NVIDIA V100 16GBnvidia16GB1no plain run of this model

ik_llama.cpp against other engines

Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.

Model and GPUik_llama.cppOther engine
Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell40.0 t/sone run, 200K contextllama.cpp105.5 t/smedian of 2, 2K to 240K context
Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell40.0 t/sone run, 200K contextNInfer172.0 t/sone run, 512 context

ik_llama.cpp results by GPU

Every card people have run ik_llama.cpp on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.

ik_llama.cpp on NVIDIA RTX 30902 reports

ik_llama.cpp on NVIDIA RTX Pro 6000 Blackwell2 reports

ik_llama.cpp on NVIDIA RTX 3060 12GB1 report

ik_llama.cpp on NVIDIA RTX 50901 report

ik_llama.cpp on NVIDIA T4 16GB1 report

ik_llama.cpp, GPU not identified1 report

Top models on ik_llama.cpp

Frequently asked

Is ik_llama.cpp faster than other engines?

These counts come from reports with the same card, model size and bit class on both engines, plain runs only. Against llama.cpp there is one matched comparison, and llama.cpp is faster. Against NInfer there is one matched comparison, and NInfer is faster. Context length and engine build still differ between the runs, so each comparison is a single data point.