llamaperf

NInfer

An inference engine for running open-weight LLMs locally.

34 community reports

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running NInfer

GPUVRAMReportsMedian t/s, Qwen3.8 125B · 6B active 4-bit
NVIDIA RTX 5090nvidia32GB17no plain run of this model
NVIDIA RTX 4090nvidia24GB6no plain run of this model
NVIDIA RTX Pro 6000 Blackwellnvidia96GB3172.0
NVIDIA RTX 3090nvidia24GB2no plain run of this model
NVIDIA RTX 5080nvidia16GB2no plain run of this model
NVIDIA RTX PRO 6000 Max-Qnvidia96GB2no plain run of this model
NVIDIA RTX 2080 Ti 22GB (modded)nvidia22GB1no plain run of this model
NVIDIA RTX 5060 Ti 16GBnvidia16GB1no plain run of this model

NInfer against other engines

Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.

Model and GPUNInferOther engine
Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell172.0 t/sone run, 512 contextik_llama.cpp40.0 t/sone run, 200K context
Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell172.0 t/sone run, 512 contextllama.cpp105.5 t/smedian of 2, 2K to 240K context

NInfer results by GPU

Every card people have run NInfer on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.

NInfer on NVIDIA RTX 509014 reports

All 14 NInfer reports on the NVIDIA RTX 5090

NInfer on NVIDIA RTX 40905 reports

NInfer on NVIDIA RTX 30902 reports

NInfer on NVIDIA RTX 50802 reports

NInfer on NVIDIA RTX 5060 Ti 16GB1 report

NInfer on NVIDIA RTX 5070 Ti1 report

NInfer on NVIDIA RTX Pro 4000 Blackwell1 report

NInfer, GPU not identified1 report

Top models on NInfer

Frequently asked

Is NInfer faster than other engines?

These counts come from reports with the same card, model size and bit class on both engines, plain runs only. Against ik_llama.cpp there is one matched comparison, and NInfer is faster. Against llama.cpp there is one matched comparison, and NInfer is faster. Context length and engine build still differ between the runs, so each comparison is a single data point.