TabbyAPI
An inference engine for running open-weight LLMs locally.
1 community report
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running TabbyAPI
| GPU | VRAM | Reports | Fastest t/s |
|---|---|---|---|
| RTX 3090nvidia | 24GB | 1 | 123.0 |