llamaperf

NVIDIA RTX 4060 Laptop 8GB

NVIDIA · 8GB · 2 reports

Engines people use on it: FreeToken 1 · llama.cpp 1

Run models on your NVIDIA RTX 4060 Laptop 8GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 4060 Laptop 8GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 8 GB of VRAM.

This page is thin (2 of 3 reports needed for indexing). Help fill it in.

Cyber-Tiel 35B (3B active) Coder

NVIDIA RTX 4060 Laptop 8GB · llama.cpp · 131,072 ctx

reported speed:
23.0 tokens/s generation · 35.0 tokens/s prompt processing
quant:
UD-IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Cyber-Tiel Coder 35B-A3B at 23 t/s decode and 35 t/s prefill on an RTX 4060 Laptop 8GB with 16 GB system RAM, at 131,072 context. Setup is llama.cpp with the UD-IQ3_XXS GGUF, q4_0 KV cache, flash attention, 30 MoE layers offloaded to CPU, and MTP speculative decoding with 2 draft tokens. Figures are at 4K context and depend on free RAM for mmap caching; a 13-prompt benchmark table is promised later.

Sep 23, 2026
Tone: positive
reported speed:
39.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports Qwen3.6 35B at 39 t/s on an 8 GB RTX 4060 laptop using the FreeToken engine. User also reports DeepSeek-V4-Flash 284B at 22-25 t/s on an RTX 5090 and GLM-5.2 753B at 15 t/s on an RTX PRO 6000.

Sep 7, 2026

Get a weekly email of new NVIDIA RTX 4060 Laptop 8GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Cyber-Tiel 35B (3B active) Coder
NVIDIA RTX 4060 Laptop 8GB
UD-IQ3_XXS
llama.cpp
131,07223.0 tokens/s
Qwen3.6 35B (3B active)
NVIDIA RTX 4060 Laptop 8GB
Not reported
FreeToken
Not reported39.0 tokens/s