llamaperf

TabbyAPI

An inference engine for running open-weight LLMs locally.

1 community report

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running TabbyAPI

GPUVRAMReportsFastest t/s
RTX 3090nvidia24GB1123.0

Top models on TabbyAPI