Nemotron-3-Ultra 550B (55B active)
2× NVIDIA DGX Spark · llama.cpp
- reported speed:
- 5.2 tokens/s generation · 120.0 tokens/s prompt processing
- quant:
- UD-Q2_K_XL (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Nemotron-3-Ultra-550B-A55B at ~5.2 tok/s decode and ~120 tok/s prefill across 2x DGX Spark (GB10, 128GB unified each) linked over 200GbE ConnectX-7 RoCE. Setup is llama.cpp upstream master with UD-Q2_K_XL GGUF (~188 GiB, 6 shards), layer-split via RPC with --tensor-split 1,1, --no-mmap and -fit off, ~95 GiB per node. User notes decode is RPC-round-trip-bound and slower per-token than a single node; the model only fits across two boxes. A 600-token reasoning sample took 117s.