ExLlamaV3
An inference engine for running open-weight LLMs locally.
2 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running ExLlamaV3
| GPU | VRAM | Reports | Fastest t/s |
|---|---|---|---|
| RTX 4080nvidia | 16GB | 1 | 56.5 |
| RTX 3080 20GBnvidia | 20GB | 1 | 25.0 |