Paddock
An inference engine for running open-weight LLMs locally.
2 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running Paddock
| GPU | VRAM | Reports | Median t/s |
|---|---|---|---|
| NVIDIA RTX Pro 6000 Blackwellnvidia | 96GB | 2 | no plain run of this model |
Paddock against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
No matched pair yet. No card has plain runs of one model size at one bit class on Paddock and on another engine, so llamaperf can't say how it compares on speed. Add a run.
Top models on Paddock
Frequently asked
Is Paddock faster than other engines?
llamaperf has no matched comparison for Paddock yet: no card has plain runs of the same model size at the same bit class on Paddock and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.