llamaperf

RTX 2080 Ti

NVIDIA · 11GB · 2 reports

See what fits on this GPU →

Use the calculator to check which models fit in 11 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →
This page is thin (2 of 3 reports needed for indexing). Help fill it in.

Qwen3.8 27B

RTX 2080 Ti · NInfer · 128,000 ctx

Tone: positive
reported speed:
45.0 tokens/s generation
quant:
W8A16
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Speculative decoding with MTP3 draft window yields ~456 tok/s, ~65% acceptance rate. Standard autoregressive ~25 tok/s. VRAM usage ~17.5 GiB with draft weights, leaving ~4.5-5.0 GiB for KV cache.

Tone: positive
reported speed:
255.0 tokens/s prompt processing
quant:
W8A8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

Custom Turing CUDA kernels for W8A8 INT8 matmul. Heterogeneous inference with 4x 11/22GB VRAM and 1TB system RAM. Computation-communication overlap for MoE routing. Open-sourced on GitHub.