Nemotron 3.5 Lightning 30B (3B active)
on NVIDIA DGX Spark · Ollama · 262,144 ctx
Oct 4, 2026
Use cases
agentictool-use
Summary
User reports Nemotron 3.5 Lightning 30B-A3B at 72 to 87 tok/s single-stream decode and about 2,600 tok/s prefill on a DGX Spark (GB10, 128GB unified memory).
Setup is Ollama 0.32.9 with the Q4_K_M GGUF at the default 262,144 token context, 26GB resident at 100% GPU, with the built-in MTP speculative decoding active.
Decode depends on the workload: 72 tok/s on prose and 84 to 87 tok/s on JSON and summaries. The same model served with vLLM using the NVFP4 checkpoint and DSpark draft model reached 108 tok/s decode and about 5,400 tok/s prefill. On one agent prompt the model answered in 485 tokens and 5.9s against 1,953 tokens and 26.0s for qwen3.5:35b-a3b.