llamaperf

Local model performance on your hardware

Find which open-weight LLMs fit in your GPU or Mac and compare the speeds people report on setups like yours.

What runs on your hardware?

Pick your GPU or Mac, then read the speeds people reported on it, or estimate which models fit and how fast they run.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Local model performance reports from the community

These are individual setups, not a controlled benchmark. Compare GPU count, quantization, context and offloading before comparing speeds. How to read a report →

Model: Llama 3
Compare setup details

Exact recorded values. Context may be a configured limit; matching filters does not establish identical prompts, offloading or concurrency.

Llama 3 8B

2× NVIDIA RTX 3090 · llama.cpp · 65,536 ctx

Tone: negative
reported speed:
129.1 tokens/s generation · 5124.4 tokens/s prompt processing
quant:
Q4_0 (GGUF)
kv:
f16
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks the new tensor parallel (split mode tensor) implementation in llama.cpp against graph parallel in ik_llama.cpp on a 2x RTX 3090 system with 48 GiB of VRAM. Setup is llama.cpp with Q4_0 quantized Llama 3 8B, f16 KV cache, 65536 context, 100 GPU layers, and flash attention enabled. The run crashed with CUDA out of memory at 34816 tokens of context. User reports 129.08 t/s generation and 5124.38 t/s prompt processing at zero context, degrading to 26.89 t/s generation and 3092.68 t/s prompt processing at 34816 tokens. User calls the PR a gimmick not ready for prime time, noting memory is not released and performance lags well behind ik_llama.cpp graph parallel.

Oct 8, 2026
Showing 1–1 of 1
Page 1 of 1

Community benchmarks snapshot

Records by GPU

NVIDIA RTX 3090167NVIDIA RTX 5090114AMD Strix Halo 128GB83NVIDIA DGX Spark65NVIDIA RTX 5060 Ti 16GB53NVIDIA RTX Pro 6000 Blackwell51NVIDIA RTX 3060 12GB46AMD Radeon AI PRO R9700 32GB43NVIDIA RTX 409037NVIDIA RTX 5070 Ti31

Records by model

1423 total
Qwen3.8788
Qwen3.6172
DeepSeek V4 Flash127
Gemma 462
Qwen3.532
DeepSeek V4.1 Flash23
other219

Records by engine

1095 total
llama.cpp574
vLLM160
Strata48
NInfer40
Ollama34
other239

Use cases

coding 463agentic 295long-context 217tool-use 124vision 89summarization 47math 37creative-writing 31multilingual 19text-generation 9rp 6reasoning 3
coding463agentic295long-context217tool-use124vision89summarization47math37creative-writing31

Median t/s by GPU

On Qwen3.8 27B at 4-bit, plain single-GPU runs. Full ranking

RTX 509093RTX Pro 600067M5 Ultra 256GB50RX 7900 XTX41RTX 309036V100 32GB33RTX 409032RX 7800 XT 16GB30RTX 5090 Laptop 24GB30Radeon AI PRO R9700 32GB29

Reports by model size

Qwen3.8 27B473Qwen3.8 125B · 6B active285DeepSeek V4 Flash 284B · 13B active102Qwen3.6 35B · 3B active101Qwen3.6 27B69Gemma 4 26B · 4B active26DeepSeek V4.1 Flash 552B · 16B active23GLM-5.3 320B · 18B active18

Quants

Q4_K_M120NVFP499IQ4_XS60Q4_K_XL57UD-Q4_K_XL50IQ3_XXS38Q437Q8_0354-bit26MXFP424