- generation:
- 14.4 tokens/s
- quant:
- Q4_K (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.
User reports Nemotron-3-Super 120B at ~14.4 t/s on a DGX Spark GB10 with 128GB unified memory. Setup is llama.cpp built natively for sm_121 (commit 463b6a963, CUDA 13.0, driver 580.126.09) with the ggml-org Q4_K GGUF (66GB). Ollama's Q4_K_M build reached ~14.2 t/s but its MoE GGUF blobs are incompatible with upstream llama.cpp (blk.1.ffn_down_exps.weight shape mismatch, expected 4096 got 1024). The ggml-org Q4_K file saves about 20GB versus Ollama's 86GB. User notes an OOM pitfall on load that requires dropping the page cache first.