llamaperf

NVIDIA RTX 5060 Laptop 8GB

NVIDIA · 8GB · 3 reports

Engines people use on it: llama.cpp 2

Run models on your NVIDIA RTX 5060 Laptop 8GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 5060 Laptop 8GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 8 GB of VRAM.

reported speed:
29.3 tokens/s generation
quant:
PTQ1_0 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

math

User reports Ternary Bonsai 2 27B at 29.27 t/s on an RTX 5060 Laptop 8 GB under Windows, using a Prism llama.cpp build. Setup is llama.cpp with PTQ1_0 (5.53 GiB), temp 0, seed 42, a 2048 reasoning budget and a 3072 token cap on 50 MATH-500 questions. A −2 logit bias on "wait", "maybe" and "perhaps" (nine token ids) dropped accuracy from 44/50 to 43/50 and raised average tokens from 845.2 to 872.3; the biased run measured 29.28 t/s. The user calls it one deterministic data point, not proof.

Oct 3, 2026
Tone: positive
reported speed:
300-400 tokens/s prompt processing

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Ornith 1.5 35B-A3B at 30-40 t/s generation and 300-400 t/s prompt processing on an RTX 5060 Laptop 8GB with 32GB DDR5 RAM. Setup uses llama.cpp via OpenCode, with generation at 30 t/s for 100k-140k context and up to 40 t/s below 100k context. User also ran Qwen 3.8 27B on the same hardware and compares the two models for coding tasks, noting Ornith handles contained tasks well but struggles with complex multi-file changes.

Sep 26, 2026

Gemma 4 12B

NVIDIA RTX 5060 Laptop 8GB · llama.cpp · 4,096 ctx

reported speed:
40.0 tokens/s generation · 700-900 tokens/s prompt processing
quant:
Q6_K (GGUF)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Gemma 4 12B at 40.0 t/s on an RTX 5060 Laptop GPU with 8 GB VRAM, using MTP speculative decoding with a Q6_K draft head and F16 KV cache at 4096 context. Setup is llama.cpp with mainline draft-mtp support (commit b9193 or later), an imatrix-guided per-tensor fit-to-VRAM quant at 6.18 GB, 48 layers, parallel 1. Prefill is 700 to 900 t/s. Without MTP the same setup reaches 27.4 t/s. With a Q8 KV cache it reaches 40.4 t/s. In real chat use the user sees 42 or more t/s on fresh context, dropping to 33 to 37 t/s with 16k of context filled.

Sep 24, 2026

Get a weekly email of new NVIDIA RTX 5060 Laptop 8GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Bonsai 2 27B
NVIDIA RTX 5060 Laptop 8GB
PTQ1_0
llama.cpp
Not reported29.3 tokens/s
Gemma 4 12B
NVIDIA RTX 5060 Laptop 8GB
Q6_K
llama.cpp
4,09640.0 tokens/s