llamaperf

Nemotron 3.5 Lightning 30B (3B active)

on NVIDIA DGX Spark · Ollama · 262,144 ctx

Tone: positive
Oct 4, 2026
Throughput
72-87 t/s gen
Quant
Q4_K_M (GGUF)
System RAM
128 GB

Use cases

agentictool-use

Summary

User reports Nemotron 3.5 Lightning 30B-A3B at 72 to 87 tok/s single-stream decode and about 2,600 tok/s prefill on a DGX Spark (GB10, 128GB unified memory). Setup is Ollama 0.32.9 with the Q4_K_M GGUF at the default 262,144 token context, 26GB resident at 100% GPU, with the built-in MTP speculative decoding active. Decode depends on the workload: 72 tok/s on prose and 84 to 87 tok/s on JSON and summaries. The same model served with vLLM using the NVFP4 checkpoint and DSpark draft model reached 108 tok/s decode and about 5,400 tok/s prefill. On one agent prompt the model answered in 485 tokens and 5.9s against 1,953 tokens and 26.0s for qwen3.5:35b-a3b.