SGLang
An inference engine for running open-weight LLMs locally.
24 community reports
This engine doesn't yet have an editorial profile on llamaperf. The community reports below show how it's been used in practice across different hardware.
Top GPUs running SGLang
| GPU | VRAM | Reports | Median t/s, Qwen3.8 27B 2-bit |
|---|---|---|---|
| NVIDIA RTX Pro 6000 Blackwellnvidia | 96GB | 6 | no plain run of this model |
| NVIDIA DGX Sparknvidia | 128GB | 5 | no plain run of this model |
| NVIDIA RTX 5090nvidia | 32GB | 5 | no plain run of this model |
| NVIDIA RTX 3090nvidia | 24GB | 3 | no plain run of this model |
| NVIDIA RTX 4090nvidia | 24GB | 2 | 67.0 |
| AMD Instinct MI300X 192GBamd | 192GB | 1 | no plain run of this model |
| NVIDIA H100 80GBnvidia | 80GB | 1 | no plain run of this model |
| NVIDIA H200nvidia | 141GB | 1 | no plain run of this model |
SGLang against other engines
Pairs of plain runs on the same card, of the same model size at the same bit class: one device, one request, no speculative decoding, the whole model in memory. Context length and build still differ between the two sides, and each side shows its own.
| Model and GPU | SGLang | Other engine |
|---|---|---|
| Qwen3.8 125B · 6B active 4-bitNVIDIA DGX Spark | 35.0 t/sone run, 256K context | llama.cpp24.5 t/sone run, 128K context |
SGLang results by GPU
Every card people have run SGLang on, with each report's model, quant and speed, newest first. Runs on several cards, with speculative decoding, with batched requests or with part of the model in system RAM say so, since each describes a different setup.
SGLang on NVIDIA RTX Pro 6000 Blackwell6 reports
- DeepSeek V4.1 Flash 552B · 16B active · FP4200.9 t/s, 1,943 t/s prompt(4 cards, speculative, part in system RAM)
- Qwen3.8 27B · NVFP4210.0 t/s(speculative)
- GLM-5.3 320B · 18B active · NVFP463.7 t/s(4 cards)
- Qwen3.8 27B · BF1677.6 t/s(speculative)
- Qwen3.8 125B · 6B active · NVFP411,352 t/s prompt
SGLang on NVIDIA RTX 50905 reports
- Qwen3.8 27B · NVFP4152.0 t/s(speculative)
- Qwen3.8 27B · NVFP4100.0 t/s(speculative)
- Qwen3.8 27B · NVFP4200.0 t/s
- Qwen3.8 27B · W282.6 t/s
- Qwen3.8 27B · NVFP4144.0 t/s(speculative)
SGLang on NVIDIA DGX Spark4 reports
- Qwen3.8 27B · NVFP441.0 t/s(speculative)
- Qwen3.8 27B · NVFP434.0 t/s(speculative)
- Qwen3.8 125B · 6B active · NVFP435.0 t/s
- Qwen3.8 125B · 6B active · NVFP451.1 t/s(4 cards, speculative)
SGLang on NVIDIA RTX 30903 reports
- Qwen3.8 27B · EXL3 3.0 bpw96.2 t/s(speculative)
- Qwen3.8 27B · AWQ-INT4136.8 t/s, 1,581 t/s prompt(4 cards, speculative)
- Qwen3.8 27B · 3.00bpw98.2 t/s, 914 t/s prompt(speculative)
SGLang on NVIDIA RTX 40902 reports
- Qwen3.8 27B · escha 2-bit67.0 t/s
- Qwen3.8 27B · EXL3 3bpw135.4 t/s
SGLang on NVIDIA H100 80GB1 report
- GLM-4.7 30B168.0 t/s(speculative)
SGLang on NVIDIA RTX 4080 Super1 report
- Qwen3.8 27B · 2.469 bpw59.0 t/s
SGLang on NVIDIA RTX Pro 4000 Blackwell1 report
- Qwen2.5 32B · AWQ71.4 t/s(2 cards, speculative)
SGLang on NVIDIA RTX PRO 6000 Max-Q1 report
- Qwen3.8 125B · 6B active · NVFP4240.0 t/s(speculative)
SGLang on NVIDIA V100 32GB1 report
- Qwen3.8 125B · 6B active · NVFP460.0 t/s(4 cards, part in system RAM)
SGLang, GPU not identified2 reports
- DeepSeek V4 Flash 284B · 13B active · NVFP44.8 t/s(part in system RAM)
- Qwen3.8 27B130.0 t/s, 4,000 t/s prompt(5 requests at once)
Top models on SGLang
Frequently asked
Is SGLang faster than other engines?
These counts come from reports with the same card, model size and bit class on both engines, plain runs only. Against llama.cpp there is one matched comparison, and SGLang is faster. Context length and engine build still differ between the runs, so each comparison is a single data point.