llamaperf

Nemotron-3-Ultra

NVIDIA · 1 report

Nemotron-3-Ultra VRAM requirements by size and quant →

How does Nemotron-3-Ultra run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Nemotron-3-Ultra on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Nemotron-3-Ultra yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Nemotron-3-Ultra

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.
Tone: positive
reported speed:
5.2 tokens/s generation · 120.0 tokens/s prompt processing
quant:
UD-Q2_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Nemotron-3-Ultra-550B-A55B at ~5.2 tok/s decode and ~120 tok/s prefill across 2x DGX Spark (GB10, 128GB unified each) linked over 200GbE ConnectX-7 RoCE. Setup is llama.cpp upstream master with UD-Q2_K_XL GGUF (~188 GiB, 6 shards), layer-split via RPC with --tensor-split 1,1, --no-mmap and -fit off, ~95 GiB per node. User notes decode is RPC-round-trip-bound and slower per-token than a single node; the model only fits across two boxes. A 600-token reasoning sample took 117s.

Sep 28, 2026

Get a weekly email of new Nemotron-3-Ultra reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Nemotron-3-Ultra 550B (55B active)
2× NVIDIA DGX Spark
UD-Q2_K_XL
llama.cpp
Not reported5.2 tokens/s