NInfer
An inference engine for running open-weight LLMs locally.
34 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running NInfer
| GPU | VRAM | Reports | Median t/s, Qwen3.8 125B · 6B active 4-bit |
|---|---|---|---|
| NVIDIA RTX 5090nvidia | 32GB | 17 | no plain run of this model |
| NVIDIA RTX 4090nvidia | 24GB | 6 | no plain run of this model |
| NVIDIA RTX Pro 6000 Blackwellnvidia | 96GB | 3 | 172.0 |
| NVIDIA RTX 3090nvidia | 24GB | 2 | no plain run of this model |
| NVIDIA RTX 5080nvidia | 16GB | 2 | no plain run of this model |
| NVIDIA RTX PRO 6000 Max-Qnvidia | 96GB | 2 | no plain run of this model |
| NVIDIA RTX 2080 Ti 22GB (modded)nvidia | 22GB | 1 | no plain run of this model |
| NVIDIA RTX 5060 Ti 16GBnvidia | 16GB | 1 | no plain run of this model |
NInfer against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
| Model and GPU | NInfer | Other engine |
|---|---|---|
| Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell | 172.0 t/sone run, 512 context | ik_llama.cpp40.0 t/sone run, 200K context |
| Qwen3.8 125B · 6B active 4-bitNVIDIA RTX Pro 6000 Blackwell | 172.0 t/sone run, 512 context | llama.cpp105.5 t/smedian of 2, 2K to 240K context |
NInfer results by GPU
Every card people have run NInfer on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
NInfer on NVIDIA RTX 509014 reports
- Qwen3.8 27B · NVFP4287.0 t/s, 13,700 t/s prompt(speculative)
- Qwen3.8 27B · NVFP4158.0 t/s, 7,265 t/s prompt(speculative)
- Qwen3.8 27B · NVFP4160.8 t/s, 3,269 t/s prompt(speculative)
- Qwen3.8 27B · NVFP4175.0 t/s(speculative)
- Qwen3.8 27B · NVFP4178.0 t/s(speculative)
NInfer on NVIDIA RTX 40905 reports
- Qwen3.8 27B · INT8149.0 t/s(speculative, part in system RAM)
- Qwen3.8 27B · Q4/Q5107.0 t/s, 5,008 t/s prompt(speculative)
- Bonsai 2 27B · Q4/Q5188.0 t/s(speculative)
- Qwen3.8 27B · E8 4-bit148.6 t/s, 1,849 t/s prompt(speculative)
- Qwen3.8 27B80.0 t/s, 1,200 t/s prompt
NInfer on NVIDIA RTX Pro 6000 Blackwell3 reports
- Qwen3.8 125B · 6B active · NVFP4172.0 t/s, 5,905 t/s prompt
- Qwen3.6 35B · 3B active600.0 t/s
- Qwen3.8 27B · groupwise-int104.8 t/s, 12,400 t/s prompt
NInfer on NVIDIA RTX 30902 reports
- Qwen3.8 27B71.0 t/s, 862 t/s prompt(speculative)
- Qwen3.8 27B70.2 t/s(speculative)
NInfer on NVIDIA RTX 50802 reports
- Qwen3.8 27B96.5 t/s(speculative)
- Qwen3.8 27B · Q3G64_F16S/Q4G64_F16S/Q5G64_F16S71.6 t/s, 1,381 t/s prompt(speculative)
NInfer on NVIDIA RTX 2080 Ti 22GB (modded)1 report
- Qwen3.8 27B · W8A1625.0 t/s(part in system RAM)
NInfer on NVIDIA RTX 5060 Ti 16GB1 report
- Qwen3.8 27B · NVFP476.7 t/s(2 cards, speculative)
NInfer on NVIDIA RTX 5070 Ti1 report
- Qwen3.8 27B · NVFP4220.0 t/s, 4,956 t/s prompt(2 cards, speculative)
NInfer on NVIDIA RTX Pro 4000 Blackwell1 report
- Qwen3.8 27B67.0 t/s, 785 t/s prompt(speculative)
NInfer on NVIDIA RTX PRO 6000 Max-Q1 report
- Qwen3.8 27B · groupwise-int96.8 t/s(speculative)
NInfer on NVIDIA V100 32GB1 report
- Qwen3.8 27B · NVFP4218.0 t/s(speculative)
NInfer, GPU not identified1 report
- Qwen3.6 35B · 3B active210.0 t/s, 4,000 t/s prompt(speculative)
Top models on NInfer
Frequently asked
Is NInfer faster than other engines?
These counts come from reports with the same card, model size and bit class on both engines, plain runs only. Against ik_llama.cpp there is one matched comparison, and NInfer is faster. Against llama.cpp there is one matched comparison, and NInfer is faster. Context length and engine build still differ between the runs, so each comparison is a single data point.