llamaperf
Oct 11, 2026
Throughput
19.9 t/s gen · 768.8 tokens/s prompt processing (prefill)Prompt processing measures input prefill speed; it does not tell you the time to first token.
Quant
Q4_K (GGUF)
Flash Attention
on

Summary

User benchmarks Nemotron-3-Super-120B-A12B at 19.94 t/s generation and 768.84 t/s prompt processing on an NVIDIA DGX Spark. Setup is llama.cpp with Q4_K GGUF, 65.10 GiB model size, flash attention enabled, 99 GPU layers, n_ubatch 2048. Batched benchmark shows aggregate throughput up to 56.16 t/s generation and 771.69 t/s prompt at 32 concurrent requests. A second model, Nemotron-3-Nano-4B Q8_0, was also benchmarked at 52.85 t/s generation and 2761.90 t/s prompt processing.