Nemotron 3 120B (12B active) Super
on NVIDIA DGX Spark · llama.cpp
Oct 11, 2026
Summary
User reports Nemotron-3-Super 120B at ~14.4 t/s on a DGX Spark GB10 with 128GB unified memory.
Setup is llama.cpp built natively for sm_121 (commit 463b6a963, CUDA 13.0, driver 580.126.09) with the ggml-org Q4_K GGUF (66GB).
Ollama's Q4_K_M build reached ~14.2 t/s but its MoE GGUF blobs are incompatible with upstream llama.cpp (blk.1.ffn_down_exps.weight shape mismatch, expected 4096 got 1024). The ggml-org Q4_K file saves about 20GB versus Ollama's 86GB. User notes an OOM pitfall on load that requires dropping the page cache first.