llamaperf

Llama 3

1 report

Llama 3 VRAM requirements by size and quant →

How does Llama 3 run on your hardware?

Pick your GPU or Mac, then read the speeds people reported for Llama 3 on it, or estimate memory fit and speed.

Free to use. No account needed. Memory estimates and community measurements are labelled separately.

Run Llama 3 yourself? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance for Llama 3

Filter this model’s reports by setup →
Thin page (1 of 3 reports needed for indexing). Add yours.

Llama 3 8B

2× NVIDIA RTX 3090 · llama.cpp · 65,536 ctx

Tone: negative
reported speed:
129.1 tokens/s generation · 5124.4 tokens/s prompt processing
quant:
Q4_0 (GGUF)
kv:
f16
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User benchmarks the new tensor parallel (split mode tensor) implementation in llama.cpp against graph parallel in ik_llama.cpp on a 2x RTX 3090 system with 48 GiB of VRAM. Setup is llama.cpp with Q4_0 quantized Llama 3 8B, f16 KV cache, 65536 context, 100 GPU layers, and flash attention enabled. The run crashed with CUDA out of memory at 34816 tokens of context. User reports 129.08 t/s generation and 5124.38 t/s prompt processing at zero context, degrading to 26.89 t/s generation and 3092.68 t/s prompt processing at 34816 tokens. User calls the PR a gimmick not ready for prime time, noting memory is not released and performance lags well behind ik_llama.cpp graph parallel.

Oct 8, 2026

Get a weekly email of new Llama 3 reports on any GPU.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Llama 3 8B
2× NVIDIA RTX 3090
Q4_0
llama.cpp
65,536129.1 tokens/s