llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Looking for a particular model?

Search for a model to see its reported speeds across GPUs and Macs.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Nemotron 3
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

NVIDIA hardware
generation:
14.4 tokens/s
quant:
Q4_K (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.

User reports Nemotron-3-Super 120B at ~14.4 t/s on a DGX Spark GB10 with 128GB unified memory. Setup is llama.cpp built natively for sm_121 (commit 463b6a963, CUDA 13.0, driver 580.126.09) with the ggml-org Q4_K GGUF (66GB). Ollama's Q4_K_M build reached ~14.2 t/s but its MoE GGUF blobs are incompatible with upstream llama.cpp (blk.1.ffn_down_exps.weight shape mismatch, expected 4096 got 1024). The ggml-org Q4_K file saves about 20GB versus Ollama's 86GB. User notes an OOM pitfall on load that requires dropping the page cache first.

Oct 11, 2026
NVIDIA hardware
generation:
19.9 tokens/s
prompt processing (prefill):
768.8 tokens/s
quant:
Q4_K (GGUF)
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.

User benchmarks Nemotron-3-Super-120B-A12B at 19.94 t/s generation and 768.84 t/s prompt processing on an NVIDIA DGX Spark. Setup is llama.cpp with Q4_K GGUF, 65.10 GiB model size, flash attention enabled, 99 GPU layers, n_ubatch 2048. Batched benchmark shows aggregate throughput up to 56.16 t/s generation and 771.69 t/s prompt at 32 concurrent requests. A second model, Nemotron-3-Nano-4B Q8_0, was also benchmarked at 52.85 t/s generation and 2761.90 t/s prompt processing.

Oct 11, 2026
Tone: positiveNVIDIA hardware
generation:
43.0 tokens/s
quant:
~3 bits per weight

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.

User reports Nemotron 3 Super (120B, 12B active) at 43 tok/s on a single RTX 4090. Setup uses the glyd engine with a ~3 bits per weight quant (48.7 GB), hot experts on GPU and the rest computed on CPU from RAM; on 32 GB machines the remainder streams from SSD. Same 4090 with llama.cpp and Unsloth Q2_K_XL does 17 tok/s. Also reports 37 tok/s on RTX 3090, 35 tok/s on a 16 GB card, and 13-18 tok/s on a 32 GB RAM PC. Cites GSM8K 97% and MMLU-Pro 77%.

Oct 11, 2026
Showing 1–3 of 3
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090185NVIDIA RTX 5090132AMD Strix Halo 128GB120NVIDIA DGX Spark74NVIDIA RTX Pro 6000 Blackwell59NVIDIA RTX 5060 Ti 16GB57NVIDIA RTX 3060 12GB50AMD Radeon AI PRO R9700 32GB50NVIDIA RTX 409042NVIDIA RTX 5070 Ti35

Records by model

1681 total
Qwen3.8926
Qwen3.6194
DeepSeek V4 Flash138
Gemma 473
Qwen3.546
GLM-5.335
other269

Records by engine

1279 total
llama.cpp651
vLLM182
Strata66
NInfer44
oMLX41
other295

Use cases

coding 534agentic 348long-context 258tool-use 148vision 105summarization 53math 42creative-writing 36multilingual 22text-generation 9rp 8reasoning 3
coding534agentic348long-context258tool-use148vision105summarization53math42creative-writing36

Median generation t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509082RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX48RTX 309036RTX 5090 Laptop 24GB34RTX 409032V100 32GB31Radeon AI PRO R9700 32GB30RX 7800 XT 16GB30

Prompt processing, same model

M5 Ultra 256GB1800RX 7900 XTX530RTX 3090561RTX 40901949V100 32GB941RTX 3080 20GB997Arc Pro B70512M1 Ultra 128GB153Arc Pro B60 24GB380M1 Max 32GB82

Reports by model size

Qwen3.8 27B538Qwen3.8 125B · 6B active356DeepSeek V4 Flash 284B · 13B active129Qwen3.6 35B · 3B active121Qwen3.6 27B71GLM-5.3 320B · 18B active30Gemma 4 26B · 4B active30DeepSeek V4.1 Flash 552B · 16B active27

Quants

Q4_K_M146NVFP4116Q4_K_XL66UD-Q4_K_XL63IQ4_XS63IQ3_XXS45Q442Q8_0394bit33MXFP433