llamaperf
Oct 3, 2026
Throughput
91.9 t/s gen · 1760.0 t/s pp
Quant
Q4_0 (GGUF)
System RAM
64 GB
VRAM reported
24 GB

Use cases

agenticcoding

Summary

User reports Nemotron 3.5 Lightning 30B-A3B at 91.91 tok/s decode on 2x Intel Arc Pro B60 24GB. Setup is llama.cpp SYCL with Q4_0 weights and an MTP Q8_0 drafter at --spec-draft-n-max 7, 22.18 GiB VRAM, 1,760 tok/s prefill at 12K context. MTP acceptance is 99.5-100% at every n-max; the model needs the whole card with no co-residence.