Oct 11, 2026
Summary
User reports Nemotron 3 Super (120B, 12B active) at 43 tok/s on a single RTX 4090.
Setup uses the glyd engine with a ~3 bits per weight quant (48.7 GB), hot experts on GPU and the rest computed on CPU from RAM; on 32 GB machines the remainder streams from SSD.
Same 4090 with llama.cpp and Unsloth Q2_K_XL does 17 tok/s. Also reports 37 tok/s on RTX 3090, 35 tok/s on a 16 GB card, and 13-18 tok/s on a 32 GB RAM PC. Cites GSM8K 97% and MMLU-Pro 77%.