TensorSharp
An inference engine for running open-weight LLMs locally.
8 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running TensorSharp
| GPU | VRAM | Reports | Median t/s, Muse 30B 8-bit |
|---|---|---|---|
| NVIDIA A40 48GBnvidia | 48GB | 4 | no plain run of this model |
| NVIDIA A100 80GBnvidia | 80GB | 1 | no plain run of this model |
| NVIDIA RTX 2000 Adanvidia | 16GB | 1 | no plain run of this model |
| NVIDIA RTX 3080 Laptop 16GBnvidia | 16GB | 1 | no plain run of this model |
| NVIDIA RTX Pro 6000 Blackwellnvidia | 96GB | 1 | 35.0 |
TensorSharp against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
No matched pair yet. No card has plain runs of one model size at one bit class on TensorSharp and on another engine, so llamaperf can't say how it compares on speed. Add a run.
TensorSharp results by GPU
Every card people have run TensorSharp on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
TensorSharp on NVIDIA A40 48GB3 reports
- DeepSeek V4.1 Flash 552B · 16B active · Q2_K41.1 t/s(8 cards)
- DeepSeek V4.1 Flash 552B · 16B active · Q4_K_M32.5 t/s, 492 t/s prompt(8 cards, part in system RAM)
- DeepSeek V4 Flash 284B · 13B active · Q8_K_XL31.5 t/s, 836 t/s prompt(4 cards, part in system RAM)
TensorSharp on NVIDIA A100 80GB1 report
- Qwen3.8 125B · 6B active · UD-Q2_K_XL56.8 t/s, 1,019 t/s prompt(2 cards)
TensorSharp on NVIDIA RTX 2000 Ada1 report
- Gemma 4 · Q8_051.7 t/s, 2,488 t/s prompt(2 cards)
TensorSharp on NVIDIA RTX 3080 Laptop 16GB1 report
- Qwen3.8 125B · 6B active11.1 t/s(part in system RAM)
TensorSharp on NVIDIA RTX Pro 6000 Blackwell1 report
- Muse 30B · Q8_035.0 t/s
Top models on TensorSharp
Frequently asked
Is TensorSharp faster than other engines?
llamaperf has no matched comparison for TensorSharp yet: no card has plain runs of the same model size at the same bit class on TensorSharp and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.