llamaperf

NVIDIA RTX 4060

NVIDIA · 8GB · 3 reports

Engines people use on it: llama.cpp 2 (reports) · LM Studio 1 (reports)

Run models on your NVIDIA RTX 4060? Add your numbers.

Your coding agent can run a speed test on your machine and send it here with a free account. Or paste a result you've already posted.

Reported performance on the NVIDIA RTX 4060

These reports mix GPU counts and memory setups, so a model listed here may have needed extra cards or CPU offloading. The calculator estimates which models fit in 8 GB of VRAM.

Qwen3.8 27B

NVIDIA RTX 4060 · llama.cpp · 16,384 ctx

Tone: positiveNVIDIA hardware
generation:
6.0 tokens/s
quant:
UD-Q4_K_XL (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is shown only when the source reports it; generation speed cannot tell us time to first token. Check the full setup before comparing.

codingagenticvision

User reports Qwen3.8-27B at ~5.99 tok/s with MTP speculative decoding on an RTX 4060 8GB with 32GB system RAM. Setup is llama.cpp with UD-Q4_K_XL GGUF at 16K context, one parallel slot, partial CPU offload since the 17.56GB model exceeds 8GB VRAM. MTP accepted 99 of 126 draft tokens (~79%), improving decode by roughly 57% over the ~3.82 tok/s without MTP. The IQ4_XS quant measured ~4.17 tok/s. User also ran Qwen3-VL 4B Q4_K_M invoice extraction: 5 of 6 image-only invoices completed in 8-38 sec, and 4 of 6 with OCR transcripts in 25-40 sec. In a 2-invoice comparison Qwen matched 22 of 24 top-level fields versus Gemma's 5 of 24. User notes these are application-level results, not standardized benchmarks.

Oct 11, 2026
Tone: mixedNVIDIA hardware
generation:
24.2 tokens/s
prompt processing (prefill):
65.6 tokens/s
quant:
IQ4_XS (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.

User reports Qwen3.6-35B-A3B at 24.2 t/s generation and 65.62 t/s prompt eval on an RTX 4060 8GB with 32GB DDR5 RAM. Setup is llama.cpp with IQ4_XS quant and --moe-cache-mib 2048, using the PR#29887 MoE expert cache in host memory. User is not getting expected t/s and asks for help; also tried Qwen3.8-Flash-Next Q2 at 8.02 t/s generation and 6.96 t/s prompt eval.

Oct 8, 2026
Tone: positiveNVIDIA hardware
generation:
75.0 tokens/s
prompt processing (prefill):
1500.0 tokens/s
quant:
IQ (GGUF)

Reported by the source; GPU count, offloading and concurrent requests can change these figures. Prompt-processing speed is input throughput, not time to first token. Check the full setup before comparing.

User reports Gemma 4 26B at 75 t/s generation and 1500 t/s prompt processing on 2x RTX 4060 8GB. Setup is LM Studio with an IQ quant, serving to Hermes. User credits mradermacher's IQ quants for getting the model running.

Sep 30, 2026

Get a weekly email of new NVIDIA RTX 4060 reports and newly released models that fit it.

Email me new reports

Reported configurations

Individual observations from this page, not expected speeds or a ranking. Select a model to open its report, with the source and offloading conditions. Reported context may be a limit rather than actual input length.

Model / hardwareQuant / engineReported contextGeneration
Qwen3.8 27B
NVIDIA RTX 4060
UD-Q4_K_XL
llama.cpp
16,3846.0 tokens/s
Qwen3.6 35B (3B active)
NVIDIA RTX 4060
IQ4_XS
llama.cpp
Not reported24.2 tokens/s
Gemma 4 26B
2× NVIDIA RTX 4060
IQ
LM Studio
Not reported75.0 tokens/s