Nemotron 3.5 Lightning 30B (3B active)
2× Intel Arc Pro B60 24GB · llama.cpp
- reported speed:
- 91.9 tokens/s generation · 1760.0 tokens/s prompt processing
- quant:
- Q4_0 (GGUF)
Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.
User reports Nemotron 3.5 Lightning 30B-A3B at 91.91 tok/s decode on 2x Intel Arc Pro B60 24GB. Setup is llama.cpp SYCL with Q4_0 weights and an MTP Q8_0 drafter at --spec-draft-n-max 7, 22.18 GiB VRAM, 1,760 tok/s prefill at 12K context. MTP acceptance is 99.5-100% at every n-max; the model needs the whole card with no co-residence.