llamaperf

Nemotron 3.5 Lightning 30B (3B active)

on 2× NVIDIA RTX 5060 Ti 16GB · llama.cpp · 1,048,576 ctx

Tone: positive
Sep 28, 2026
Throughput
80.0 t/s gen · 4737.5 t/s pp
Quant
NVFP4 (GGUF)
KV cache
q8_0
VRAM reported
32 GB

Summary

User reports Nemotron 3.5 Lightning 30B-A3B at 79.98 t/s generation and 4737.46 t/s prompt processing on 2x RTX 5060 Ti 16GB. Setup is llama.cpp with NVFP4 GGUF weights and q8_0 KV cache, 1048576 token context, flash attention on, no MTP layers. The run processed a 54025-token prompt and generated 2156 tokens; the user notes it fits in 32 GB VRAM with no expert-layer offloading to RAM.