Nemotron 3.5 Lightning 30B (3B active)
2× NVIDIA Tesla P100 16GB · llama.cpp
- reported speed:
- 110.0 tokens/s generation
- quant:
- Q4_0 (GGUF)
- kv:
- F16
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Nemotron-3.5-Lightning-30B-A3B at 110 tok/s writing code on two Tesla P100 16GB cards. Setup is llama.cpp b10970 with Q4_0 weights, F16 KV cache, tensor split across both cards, and the model's built-in MTP draft head at n-max 2. The same model reaches 95 tok/s on prose, 50 tok/s at 128k context, 39 tok/s at 256k, and 16 tok/s at 1M tokens. The study covers 545 speed measurements of 29 models from 2B to 122B parameters, all weights and KV cache in VRAM with no system RAM offload. A 119B MoE model runs 39 tok/s against 4.3 tok/s for a 70B dense model on the same cards. Tensor split makes dense models from 8B up 21-44% faster. A single P100 throttles to 906 MHz and loses 25% under sustained load, while two cards share the heat and lose 5.5%.