llamaperf

ExLlamaV3

An inference engine for running open-weight LLMs locally.

2 community reports

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running ExLlamaV3

GPUVRAMReportsFastest t/s
RTX 4080nvidia16GB156.5
RTX 3080 20GBnvidia20GB125.0

Top models on ExLlamaV3