llamaperf

NVIDIA RTX 2080 Ti 22GB (modded)

NVIDIA · 22GB · 7 reports

As of 7 Oct 2026, the models most run on the NVIDIA RTX 2080 Ti 22GB (modded), with the median of plain runs (one device, one request, no speculative decoding):

Engines people use on it: llama.cpp 3 · NInfer 1 · vLLM 1

Run models on your NVIDIA RTX 2080 Ti 22GB (modded)? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 2080 Ti 22GB (modded)

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 22 GB of VRAM.

Tone: positive
reported speed:
69.0 tokens/s generation
quant:
UD-IQ4_XS (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.6-35B-A3B at 4-bit (UD-IQ4_XS) scoring 89.63% pass@1 on HumanEval at 69 tok/s on a single RTX 2080 Ti 22GB. Setup is llama.cpp with UD-IQ4_XS dynamic 4-bit (4.25 bpw), q8_0 KV cache, 16384 context, flash-attn on, whole model resident in VRAM with no offload. A routing patch (MoE expansion, 20 experts instead of 8 on layers 25-39) reached 90.85% pass@1 at 56 tok/s, a 19% decode speed cost; the user calls the +2 problems within statistical noise and notes the tests are original HumanEval, not EvalPlus.

Oct 7, 2026
Tone: positive
reported speed:
59-68 tokens/s generation
quant:
GPTQ-Int4 (GPTQ)
kv:
turboquant_k3v4_nc

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentictool-usevisionlong-context

User reports Qwen3.8-27B at ~59-68 tok/s single-stream decode on 2x modded RTX 2080 Ti 22GB (SM75, TP=2, NVLink). Setup is a vLLM fork with GPTQ-Int4 self-quant, turboquant_k3v4_nc KV cache, 524,288 max-model-len (YaRN 2x), MTP K=2 speculative decoding, and vision enabled. Aggregate throughput is 364.7 tok/s at 16 lanes; 24 lanes regresses. The S4 scoped re-emission drafter measured 811.5 tok/s on copy-shaped spans but crashes at current HEAD, and v3 GPU-merge async is negative.

Oct 7, 2026
Tone: positive
reported speed:
95.5 tokens/s generation · 465.8 tokens/s prompt processing
quant:
Q4_K_P (GGUF)
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen3.6-35B-A3B at 95.5 t/s decode and 465.8 t/s prefill on a mixed three-GPU Turing setup of 2x CMP 50HX 10GB plus 1x RTX 2080 Ti 22GB. Setup is llama.cpp with a ported DP2A patch (PR #25834) and -fmad=false, Q4_K_P GGUF weights, Q8_0 KV cache, Flash Attention, 262144 context, MTP speculative decoding at --spec-draft-n-max 3, tensor split 1,1,2.5 with the RTX as tail stage, ubatch 448. Baseline DP4A gave 47.8 t/s decode and 372.7 t/s prefill; DP2A alone 54.0 t/s; adding -fmad=false 62.2 t/s. A repo-28k workload measured 405.4 t/s prefill and 82.3 t/s decode. An end-to-end rerun showed 91.4 t/s raw eval, 4% below the 95.5 reference. Context 368640 also works but quality beyond the trained window was not measured.

Oct 7, 2026
Tone: positive
reported speed:
36.6 tokens/s generation
quant:
IQ3_S (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextagentictool-usevision

User reports Qwen3.8-27B at 36.6 tok/s decode on a single modded RTX 2080 Ti 22GB, with 262,144 tokens of context. Setup is KVMem (retrieval-based long context) with IQ3_S weights, q8_0 KV cache, MTP speculative decoding (58.6%/61.6% acceptance) and vision, using 16,552 of 22,528 MiB VRAM. The user compares against the upstream author's RTX 5060 Ti 16GB at 31.7 tok/s, and notes the lossless KV-streaming alternative runs 30-42 tok/s in the resident window but drops to ~9.6 tok/s past 135K tokens; a 260,096-token needle-in-a-haystack test hit exactly.

Oct 7, 2026
Tone: positive
reported speed:
101.3 tokens/s generation · 1841.7 tokens/s prompt processing
quant:
AWQ (AWQ)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticlong-context

User reports Qwen3.6-27B-AWQ at 101.3 tok/s decode and 1841.7 tok/s prefill on dual modified RTX 2080 Ti 22GB cards with NVLink. Setup is vLLM 0.21.0 with AWQ Marlin, TP=2, MTP K=3, FlashInfer/FA2 attention, and FlashQLA SM70/SM75 legacy GDN prefill. The speed columns use the PP4096/TG128 repeat. The same rig reached a 735,084 token KV cache with turboquant_4bit_nc at max_model_len=262144, and passed a PP262000/TG1 gate at 785.26 tok/s prefill. A sequential 60-request Ragent6 run averaged 700.9 tok/s prefill and 35.2 tok/s generation. Gemma4 31B GPTQ reached 99.64 tok/s decode and 1655.65 tok/s prefill on the same runtime.

Oct 7, 2026
Tone: positive
reported speed:
44.1 tokens/s generation
quant:
IQ3_S (GGUF)
kv:
Q8_0
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

long-contextvision

User reports Qwen3.8-27B-GSQ-RCO at 44.07 t/s decode on a hardware-modded RTX 2080 Ti 22GB, with 38~44 t/s quoted as the overall range. Setup is a llama.cpp fork (TurboQuant 4-bit + Adaptive KV Streaming) with IQ3_S weights, 256K context (262,144 tokens), Q8_0 key cache and turbo4 value cache, a 2048 MiB GPU staging pool, and MTP speculative decoding with 2 draft tokens. Needle tests at 8K/16K/32K gave 41.23, 39.00 and 34.84 t/s decode with 81-83% draft acceptance and 100% recall at 82% depth; a sustained 256K run gave 40.55 t/s at 89.17% acceptance. Peak VRAM was 18,619 MiB and GPU power peaked at 266.4 W.

Sep 28, 2026
Tone: positive
reported speed:
25.0 tokens/s generation
quant:
W8A16
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at ~25 tok/s with standard autoregressive decoding (MTP0) on a modded RTX 2080 Ti 22GB. Setup is a Turing port of the NInfer engine with the official groupwise-int (W8A16) artifact and Q8 KV cache. With MTP3 speculative decoding (draft window = 3) the title gives 45 tok/s and the body ~456 tok/s, at ~65% acceptance rate. VRAM usage is ~17.5 GiB with the MTP draft weights loaded, leaving ~4.5-5.0 GiB free for KV cache.

Sep 7, 2026

Get a weekly email of new NVIDIA RTX 2080 Ti 22GB (modded) reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.6 35B (3B active)
NVIDIA RTX 2080 Ti 22GB (modded)
UD-IQ4_XS
llama.cpp
16,38469.0 tokens/s
Qwen3.6 35B (3B active) Uncensored-HauhauCS-Aggressive
3× NVIDIA RTX 2080 Ti 22GB (modded)
Q4_K_P
llama.cpp
262,14495.5 tokens/s
Qwen3.8 27B
NVIDIA RTX 2080 Ti 22GB (modded)
IQ3_S
KVMem
262,14436.6 tokens/s
Qwen3.6 27B
2× NVIDIA RTX 2080 Ti 22GB (modded)
AWQ
vLLM
4,096101.3 tokens/s
Qwen3.8 27B
NVIDIA RTX 2080 Ti 22GB (modded)
IQ3_S
llama.cpp
262,14444.1 tokens/s
Qwen3.8 27B
NVIDIA RTX 2080 Ti 22GB (modded)
W8A16
NInfer
128,00025.0 tokens/s