text-generation-webui
An inference engine for running open-weight LLMs locally.
3 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running text-generation-webui
| GPU | VRAM | Reports | Fastest t/s |
|---|---|---|---|
| RTX 3090nvidia | 24GB | 2 | 12.5 |
| RTX 4060 Ti 16GBnvidia | 16GB | 1 | 32.5 |