Nemotron 3.5 Lightning 30B (3B active)
on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 1,048,576 ctx
Sep 28, 2026
Summary
User reports Nemotron 3.5 Lightning 30B-A3B at 79.98 t/s generation and 4737.46 t/s prompt processing on 2x RTX 5060 Ti 16GB.
Setup is llama.cpp with NVFP4 GGUF weights and q8_0 KV cache, 1048576 token context, flash attention on, no MTP layers.
The run processed a 54025-token prompt and generated 2156 tokens; the user notes it fits in 32 GB VRAM with no expert-layer offloading to RAM.