llamaperf

SGLang

An inference engine for running open-weight LLMs locally.

24 community reports

This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.

Top GPUs running SGLang

GPUVRAMReportsMedian t/s, Qwen3.8 27B 2-bit
NVIDIA RTX Pro 6000 Blackwellnvidia96GB6no plain run of this model
NVIDIA DGX Sparknvidia128GB5no plain run of this model
NVIDIA RTX 5090nvidia32GB5no plain run of this model
NVIDIA RTX 3090nvidia24GB3no plain run of this model
NVIDIA RTX 4090nvidia24GB267.0
AMD Instinct MI300X 192GBamd192GB1no plain run of this model
NVIDIA H100 80GBnvidia80GB1no plain run of this model
NVIDIA H200nvidia141GB1no plain run of this model

SGLang against other engines

Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.

Model and GPUSGLangOther engine
Qwen3.8 125B · 6B active 4-bitNVIDIA DGX Spark35.0 t/sone run, 256K contextllama.cpp24.5 t/sone run, 128K context

SGLang results by GPU

Every card people have run SGLang on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.

SGLang on NVIDIA RTX 50905 reports

SGLang on NVIDIA RTX 30903 reports

SGLang on NVIDIA H100 80GB1 report

SGLang on NVIDIA RTX Pro 4000 Blackwell1 report

SGLang on NVIDIA V100 32GB1 report

SGLang, GPU not identified2 reports

Top models on SGLang

Frequently asked

Is SGLang faster than other engines?

These counts come from reports with the same card, model size and bit class on both engines, plain runs only. Against llama.cpp there is one matched comparison, and SGLang is faster. Context length and engine build still differ between the runs, so each comparison is a single data point.