- generation:
- 43.0 tokens/s
- quant:
- ~3 bits per weight
Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.
User reports Nemotron 3 Super (120B, 12B active) at 43 tok/s on a single RTX 4090. Setup uses the glyd engine with a ~3 bits per weight quant (48.7 GB), hot experts on GPU and the rest computed on CPU from RAM; on 32 GB machines the remainder streams from SSD. Same 4090 with llama.cpp and Unsloth Q2_K_XL does 17 tok/s. Also reports 37 tok/s on RTX 3090, 35 tok/s on a 16 GB card, and 13-18 tok/s on a 32 GB RAM PC. Cites GSM8K 97% and MMLU-Pro 77%.