llamaperf

Nemotron 3

1 report

As of 11 Oct 2026, Nemotron 3 120B · 12B active at 3-bit on the hardware it is most run on, with the median of runs that kept part of the model in system RAM (one device, one request):

Nemotron 3 VRAM requirements by size and quant →

How does Nemotron 3 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Nemotron 3 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Nemotron 3 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Nemotron 3

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: positiveNVIDIA hardware
generation:
43.0 tokens/s
quant:
~3 bits per weight

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.

User reports Nemotron 3 Super (120B, 12B active) at 43 tok/s on a single RTX 4090. Setup uses the glyd engine with a ~3 bits per weight quant (48.7 GB), hot experts on GPU and the rest computed on CPU from RAM; on 32 GB machines the remainder streams from SSD. Same 4090 with llama.cpp and Unsloth Q2_K_XL does 17 tok/s. Also reports 37 tok/s on RTX 3090, 35 tok/s on a 16 GB card, and 13-18 tok/s on a 32 GB RAM PC. Cites GSM8K 97% and MMLU-Pro 77%.

Oct 11, 2026

Get a weekly email of new Nemotron 3 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Nemotron 3 120B (12B active) Super
NVIDIA RTX 4090
~3 bits per weight
glyd
Not reported43.0 tokens/s