llamaperf

NVIDIA RTX 4060 Ti 8GB

NVIDIA · 8GB · 3 reports

Engines people use on it: LM Studio 1 · llama.cpp 1

Run models on your NVIDIA RTX 4060 Ti 8GB? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 4060 Ti 8GB

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 8 GB of VRAM.

Tone: positive
reported speed:
52-65 tokens/s generation
quant:
Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

agentic

User reports Qwen3.6-35B-A3B at 52-65 tok/s on an RTX 4060 Ti 8GB with 64 GB system RAM. Setup is llama.cpp with Q4_K_XL at 131k context, experts offloaded to system RAM and the rest in VRAM, on headless Linux. The user compares download defaults (~25 tok/s), tuned Windows (39-45 tok/s), and tuned headless Linux (52-65 tok/s). They also report Qwen3.8-Flash-Next 125B at 17-19 tok/s and Ternary Bonsai 27B at 36 tok/s.

Oct 6, 2026
Tone: mixed
reported speed:
8.0 tokens/s generation
quant:
Q2 (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 8 t/s with Qwen3.5 35B Q2 on an RTX 4060 Ti 8GB, with no CUDA device selected. Setup is llama.cpp with a Q2 GGUF quant; the user notes VRAM usage was 3000MB/8k and that the run had no CUDA device selected. The user is troubleshooting a cuBLAS crash and asks how to cleanly uninstall and reinstall llama.cpp with CUDA 12, and whether CUDA would improve generation speed.

Sep 29, 2026
Tone: positive
reported speed:
13.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a 26B-A4B Gemma 4 uncensored MoE running at about 13 tok/s on an RTX 4060 Ti 8GB, with experts offloaded to CPU. Setup is LM Studio with a Jev decision-model router that picks between a 3B, a 4B, a 12B and the 26B MoE; only one model fits in VRAM at a time, so a misroute costs a 15 to 45 second model swap. The router agreed with the user's labels 16 out of 16 on labelled prompts, with median latency 530ms and p95 about 643ms; a post-reply judge caught refusals 15 out of 15 on synthetic pairs and 5 of 6 on real replies.

Sep 18, 2026

Get a weekly email of new NVIDIA RTX 4060 Ti 8GB reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.5 35B (3B active)
NVIDIA RTX 4060 Ti 8GB
Q2
llama.cpp
Not reported8.0 tokens/s
Gemma 4 26B (4B active) Uncensored
NVIDIA RTX 4060 Ti 8GB
Not reported
LM Studio
Not reported13.0 tokens/s