llamaperf

NVIDIA RTX 2000 Ada

NVIDIA · 16GB · 1 report

Engines people use on it: TensorSharp 1

Run models on your NVIDIA RTX 2000 Ada? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 2000 Ada

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 16 GB of VRAM.

This page is thin (1 of 3 reports needed for indexing). Help fill it in.

Gemma 4

2× NVIDIA RTX 2000 Ada · TensorSharp

Tone: positive
reported speed:
51.7 tokens/s generation · 2488.0 tokens/s prompt processing
quant:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks TensorSharp multi-GPU tensor parallelism on 2x RTX 2000 Ada 16GB. Gemma 4 E4B Q8_0 is the primary model among several tested. TP=2 raises decode speed from 37.3 to 51.7 tok/s.

Sep 7, 2026

Get a weekly email of new NVIDIA RTX 2000 Ada reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Gemma 4
2× NVIDIA RTX 2000 Ada
Q8_0
TensorSharp
Not reported51.7 tokens/s