llamaperf

RTX 3060 12GB

NVIDIA · 12GB · 24 reports

See what fits on this GPU →

Use the calculator to check which models fit in 12 GB of VRAM. The reports below include different GPU counts and memory setups; a model listed here may need extra cards or CPU offloading.

How to compare these reports →

Unknown family 27B

RTX 3060 12GB · RAMDeck

Tone: positive
reported speed:
1.9 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a 27B model at about 16 GB sharded across 4 devices via RAMDeck, reaching 1.92 tok/s at roughly 25 ms latency. The devices are an old 12 GB Windows laptop as primary with 3.4 GB, a Mini PC with an RTX 3060 with 20 GB, a Mac mini with 3.7 GB, and an Android phone with 1 GB. The model family is not named. The user notes the run is slower than a prior 13B run and emphasizes feasibility over speed.

Sep 10, 2026

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 65,536 ctx

reported speed:
11.4 tokens/s generation · 337.9 tokens/s prompt processing
quant:
IQ3_XXS (GGUF)
kv:
q4_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 10-20 t/s on an RTX 3090 at 65,536 context, with a detailed log showing tg=11.44 t/s. Setup is llama.cpp with the Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf file, a q4_0 KV cache for both K and V, and MTP speculative decoding with draft acceptance 0.4125. Prompt processing runs ~334-338 t/s. The post title says IQ3 XXS while the body also mentions IQ3_S - 3.4375 bpw. More than 1 GB of VRAM is left after loading.

Sep 10, 2026
Tone: positive
reported speed:
20.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen3.8 27B at 20 t/s split between an RTX 3060 12GB and an RX 9070 XT, and at 5 t/s on a 780M iGPU with 5400MHz DDR5. The 20 t/s figure is for the two-GPU split configuration and the 5 t/s figure is for the iGPU. User prefers the 5 t/s iGPU setup for system usability.

Sep 9, 2026
Tone: mixed
reported speed:
38.4 tokens/s generation
quant:
IQ3_XXS
kv:
Q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports Qwen3.8 27B at 38.4 t/s and Qwen3.6 35B-A3B (MoE) at 55.9 t/s on an RTX 3060. Setup is llama.cpp with tuning, at 16K context on Ubuntu and 12K context on WSL2. The user notes editing speeds of about 188 t/s and is frustrated at not reaching 50-60 t/s for the 27B model.

Sep 9, 2026
Tone: positive
reported speed:
28.0 tokens/s generation · 315.0 tokens/s prompt processing
quant:
IQ2_M
kv:
Q8
mtp (multi-token prediction):
off

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

general-conversation

User reports Qwen3.8 Flash Next at 27-29 t/s generation and 290-340 t/s prefill across 4 GPUs using 52 GB of VRAM. Setup is the IQ2_M quant with context offloaded at Q8 and an n-gram cache on SSD. User reports roughly 98% clean Finnish output, but notes safety guardrails cause freezes in grey areas.

Sep 9, 2026
Tone: positive
reported speed:
14.2 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User benchmarks a 25-task coding-agent benchmark on an RTX 3060 12GB. KAT-Coder completed 21/25 tasks and won. Qwen3.8-27B scored 13/25 at 14.73 tps, Nemotron 3.5 Lightning 30B-A3B 8/25 at 29.96 tps, Tiel-Coder 35B-A3B 3/25 strict at 20.57 tps, and Ternary-Bonsai-27B 0/25 at 59.21 tps. User plans to tweak settings to reach 18 tps.

Sep 9, 2026
Tone: positive
reported speed:
12.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports running llama.cpp's ggml-rpc backend across a heterogeneous cluster pooling RAM and VRAM from an Acer laptop CPU, a Windows RTX 3060 on CUDA, and a Mac Mini on Metal. The primary API server runs on the weakest machine, and the setup uses llama.cpp's built-in benchmark script. The project is source-available under the Commons Clause.

Sep 9, 2026
Tone: positive
reported speed:
24.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen 3.6 27B at ~24 t/s on dual RTX 3060 GPUs. The user also discusses diffusion-based techniques, MOE, quants, and token authority as a wishlist for future local models.

Sep 7, 2026
reported speed:
3.5 tokens/s generation · 3.0 tokens/s prompt processing
quant:
IQ2_M

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports a dual RTX 3060 setup with RAM offloading, a Ryzen 7500F and 96 GB of 5600 RAM, writing a tetris game. Prompt eval runs at 3.0 tok/s. Generation is reported as 4.5 tok/s in the summary but 3.5 tok/s in PowerShell, and the user trusts PowerShell. LM Studio failed to offload to the second GPU, so the user used Unsloth Studio.

Sep 7, 2026
reported speed:
70.0 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen 3.6 35B-A3B at about 70 t/s on an RTX 3060. The figure is given as past experience rather than a benchmark run.

Sep 7, 2026

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
40.0 tokens/s generation
quant:
Q4_K_XL (GGUF)
kv:
q8_0

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports about 40 t/s with a Q4_K_XL quant on 2x RTX 3060 12GB. Setup is llama.cpp with flash attention and speculative decoding. The model wrote CUDA code but could not run it due to VRAM limits.

Sep 7, 2026

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 131,072 ctx

reported speed:
50.6 tokens/s generation · 573.0 tokens/s prompt processing
quant:
Q4_K_XL (GGUF)
kv:
q8_0
mtp (multi-token prediction):
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a dual RTX 3060 12GB setup with tensor split. MTP draft accept is 4/4 (100%) on probe. VRAM usage is 23.2/24.0 GiB.

Sep 7, 2026

Qwen3.8

RTX 3060 12GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
20.0 tokens/s generation
quant:
Q4_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports building a game in two prompts with Qwen 3.8 on 2x RTX 3060. Setup is llama.cpp with the UD-Q4_K_XL quant, 128k context, and a quantized cache at k5_0/v4_1. The user notes that 32 GB of system RAM caused issues after context compaction, which a 'continue' resolved.

Sep 7, 2026

DeepSeek V4 Flash 284B (13B active)

RTX 3060 12GB · llama.cpp · 368,640 ctx

Tone: positive
reported speed:
10.1 tokens/s generation · 99.4 tokens/s prompt processing
quant:
Q4_K_XL (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports DeepSeek-V4-Flash-0731 at 10.1 t/s generation and 99.4 t/s prompt processing on 4x RTX 3060 12GB at 368k context. Setup is llama.cpp build b10181 with a Q8_0 KV cache and a microbatch size of 2048, with the 144 GiB GGUF mostly in 128GB of system RAM. Runs at 376832 and 360448 context gave similar speeds.

Sep 7, 2026

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 77,000 ctx

Tone: positive
reported speed:
26.9 tokens/s generation · 299.3 tokens/s prompt processing
quant:
Q4_K_XL (gguf)
kv:
Q4

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports a llama.cpp RPC cluster with tensor-split 20,23. Setup uses MTP speculative decoding with 84.37% draft acceptance and a 96.14% prefix cache hit ratio. Max context was stress-tested at 72,712/77,000 tokens.

Sep 7, 2026

Qwen3.8 27B

RTX 3060 12GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
27.7 tokens/s generation · 503.7 tokens/s prompt processing
quant:
Q4_K_M (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingagentic

User reports Qwen3.8-27B at 503.7 t/s prefill and 27.7 t/s decode on dual RTX 3060 12GB cards. Setup is a Q4_K_M GGUF with a Q8_0 KV cache at 131k context, BF16 vision projector on CPU, xHigh reasoning, on a Ryzen 5 5500 with 64GB DDR4-3200 and GPUs limited to 130W each. With speculation the user measured about 503 t/s prefill and about 29 t/s decode.

Sep 7, 2026

Qwen3.8 125B (6B active) Flash-Next

RTX 3060 12GB · llama.cpp · 131,072 ctx

Tone: positive
reported speed:
12.0 tokens/s generation · 303.0 tokens/s prompt processing
quant:
IQ4_XS (gguf)
kv:
Q8

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports a 125B MoE model with 6B active on 2x RTX 3060 12GB in llama.cpp, with prefill improving from 36 to 303 t/s and decode at 12 t/s. Setup uses -sm layer and -ub 2048. ik_llama.cpp was also tested at 407 t/s prefill but uses ~108GB RAM.

Sep 7, 2026
Tone: mixed
reported speed:
13.5 tokens/s generation · 200.0 tokens/s prompt processing
quant:
UD-Q4_K_XL (gguf)
kv:
F16

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports Qwen 3.8 Flash Next at 12-15 t/s on an RTX 3060 12GB in a Thinkstation P520 with 256GB RAM, at 65k context. Setup is llama.cpp server with the model offloaded to GPU at --n-gpu-layers 999 and MoE experts on CPU at --n-cpu-moe 48. The user notes the model is slow and low context, but output quality is much better than the previous Qwen 3.6 35B A3B. Performance degrades significantly when the system is busy.

Sep 7, 2026
reported speed:
7.5 tokens/s generation

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 7-8 t/s with 2x3060 and RAM offloading, but 4.5-5.5 t/s with 3x3060. User asks why adding a GPU slows down generation, and notes there is no tensor parallelism in both chats. User asks about enabling ngram or mlock.

Sep 7, 2026

Qwen3.6 27B

RTX 3060 12GB · llama.cpp · 64,000 ctx

Tone: positive
reported speed:
43.3 tokens/s generation · 456.1 tokens/s prompt processing
quant:
Q4_K_S (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

User reports 43.26 t/s generation and 456 t/s prefill at 12k context on dual RTX 3060 cards. Setup is tensor parallel with MTP enabled and 64k context. Without MTP at 96k context, generation is 31 t/s. User praises the value and stability of CUDA.

May 27, 2026

Qwen3.6 27B

RTX 3060 12GB · llama.cpp · 32,000 ctx

Tone: positive
reported speed:
70.0 tokens/s generation · 780.0 tokens/s prompt processing
quant:
Q2-XS (gguf)
kv:
Q8
flash attention:
off
rating:
4/5

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

codingcreative-writingtool-usesummarizationvisionagenticmultilingual

User reports a dense model at 11.9 GB GGUF size, barely offloading to CPU. The user wants to try a higher quant, having chosen this one because of the 11.9 GB GGUF size.

May 12, 2026

Qwen3.6 35B (3B active)

RTX 3060 12GB · llama.cpp · 32,768 ctx

Tone: positive
reported speed:
46.8 tokens/s generation · 914.0 tokens/s prompt processing
quant:
IQ4_XS (gguf)
kv:
Q8
flash attention:
on

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

coding

User reports llama-bench results of about 914 t/s for pp512 and about 46.8 t/s for tg128. A practical coding profile at 32k context generates at about 43.4 t/s. MTP gave about 47.7 t/s, a 2% improvement.

May 9, 2026
Tone: positive
reported speed:
60.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

text-generation

User reports E2B at about 60 t/s on an RTX 3060 12GB. The figure is estimated from compute-market tiers rather than measured.

May 1, 2026
Tone: positive
reported speed:
45.0 tokens/s generation
quant:
Q4_K_M (gguf)

Reported by the source; GPU count, offloading and concurrent requests can change this figure. Check the full setup before comparing.

text-generation

User reports about 45 t/s on an RTX 3060 12GB. The E4B model fits easily.

May 1, 2026