llamaperf

Unsloth Studio

An inference engine for running open-weight LLMs locally.

5 community reports

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running Unsloth Studio

GPUVRAMReportsMedian t/s, Gemma 4 8B 4-bit
M1 16GBapple16GB18.2
NVIDIA RTX 3060 12GBnvidia12GB1no plain run of this model
NVIDIA RTX 5060 Ti 16GBnvidia16GB1no plain run of this model
NVIDIA RTX 5070 Ti Laptop 12GBnvidia12GB1no plain run of this model
NVIDIA RTX 5090nvidia32GB1no plain run of this model

Unsloth Studio against other engines

Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.

No matched pair yet. No card has plain runs of one model size at one bit class on Unsloth Studio and on another engine, so llamaperf can't say how it compares on speed. Add a run.

Unsloth Studio results by GPU

Every card people have run Unsloth Studio on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.

Unsloth Studio, GPU not identified1 report

Top models on Unsloth Studio

Frequently asked

Is Unsloth Studio faster than other engines?

llamaperf has no matched comparison for Unsloth Studio yet: no card has plain runs of the same model size at the same bit class on Unsloth Studio and on another engine. Speed claims about engines need that pairing, so this page doesn't make one.