llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Nemotron 3.5 Lightning
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Tone: positive
reported speed:
110.0 tokens/s generation
quant:
Q4_0 (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagenticlong-context

User reports Nemotron-3.5-Lightning-30B-A3B at 110 tok/s writing code on two Tesla P100 16GB cards. Setup is llama.cpp b10970 with Q4_0 weights, F16 KV cache, tensor split across both cards, and the model's built-in MTP draft head at n-max 2. The same model reaches 95 tok/s on prose, 50 tok/s at 128k context, 39 tok/s at 256k, and 16 tok/s at 1M tokens. The study covers 545 speed measurements of 29 models from 2B to 122B parameters, all weights and KV cache in VRAM with no system RAM offload. A 119B MoE model runs 39 tok/s against 4.3 tok/s for a 70B dense model on the same cards. Tensor split makes dense models from 8B up 21-44% faster. A single P100 throttles to 906 MHz and loses 25% under sustained load, while two cards share the heat and lose 5.5%.

Oct 6, 2026
Tone: positive
reported speed:
72-87 tokens/s generation
quant:
Q4_K_M (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentictool-use

User reports Nemotron 3.5 Lightning 30B-A3B at 72 to 87 tok/s single-stream decode and about 2,600 tok/s prefill on a DGX Spark (GB10, 128GB unified memory). Setup is Ollama 0.32.9 with the Q4_K_M GGUF at the default 262,144 token context, 26GB resident at 100% GPU, with the built-in MTP speculative decoding active. Decode depends on the workload: 72 tok/s on prose and 84 to 87 tok/s on JSON and summaries. The same model served with vLLM using the NVFP4 checkpoint and DSpark draft model reached 108 tok/s decode and about 5,400 tok/s prefill. On one agent prompt the model answered in 485 tokens and 5.9s against 1,953 tokens and 26.0s for qwen3.5:35b-a3b.

Oct 4, 2026
Tone: positive
reported speed:
91.9 tokens/s generation · 1760.0 tokens/s prompt processing
quant:
Q4_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agenticcoding

User reports Nemotron 3.5 Lightning 30B-A3B at 91.91 tok/s decode on 2x Intel Arc Pro B60 24GB. Setup is llama.cpp SYCL with Q4_0 weights and an MTP Q8_0 drafter at --spec-draft-n-max 7, 22.18 GiB VRAM, 1,760 tok/s prefill at 12K context. MTP acceptance is 99.5-100% at every n-max; the model needs the whole card with no co-residence.

Oct 3, 2026
Tone: positive
reported speed:
80.0 tokens/s generation · 4737.5 tokens/s prompt processing
quant:
NVFP4 (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Nemotron 3.5 Lightning 30B-A3B at 79.98 t/s generation and 4737.46 t/s prompt processing on 2x RTX 5060 Ti 16GB. Setup is llama.cpp with NVFP4 GGUF weights and q8_0 KV cache, 1048576 token context, flash attention on, no MTP layers. The run processed a 54025-token prompt and generated 2156 tokens; the user notes it fits in 32 GB VRAM with no expert-layer offloading to RAM.

Sep 28, 2026
Tone: positive
quant:
W4A16

User compares a W4A16 quant of Nemotron 3.5 Lightning 30B-A3B against an IQ4_XS GGUF on an RTX 3090. Setup is vLLM for W4A16 and llama.cpp for IQ4_XS. The two are near-parity in instruction following benchmarks, with about 4.5x throughput by B16. The user describes the model as fast and reliable, and suitable for batch labelling and agentic responses.

Sep 7, 2026
Showing 1–5 of 5
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090165NVIDIA RTX 5090111AMD Strix Halo 128GB82NVIDIA DGX Spark57NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409035NVIDIA RTX 5070 Ti30

Records by model

1385 total
Qwen3.8771
Qwen3.6170
DeepSeek V4 Flash121
Gemma 461
Qwen3.528
Qwen322
other212

Records by engine

1063 total
llama.cpp568
vLLM152
Strata46
NInfer38
Ollama34
other225

Use cases

coding 444agentic 282long-context 204tool-use 119vision 82summarization 45math 35creative-writing 30multilingual 19text-generation 9rp 6reasoning 3
coding444agentic282long-context204tool-use119vision82summarization45math35creative-writing30

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509081RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B464Qwen3.8 125B · 6B active278DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active100Qwen3.6 27B68Gemma 4 26B · 4B active25DeepSeek V4.1 Flash 552B · 16B active22Muse 30B18

Quants

Q4_K_M119NVFP493IQ4_XS56Q4_K_XL56UD-Q4_K_XL49IQ3_XXS38Q436Q8_0354-bit24Q6_K23