llamaperf

Best local LLMs for 12GB VRAM

How far 12GB of graphics memory goes for local models on cards like NVIDIA RTX 3060 12GB, NVIDIA RTX 4070, NVIDIA RTX 5070 or NVIDIA RTX 5070 Ti Laptop 12GB: which models fit, and what people actually run on it.

Is 12GB of VRAM enough?

12GB holds a dense model of up to about 12B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.5 13B, which needs about 9.4 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top. People do push further: the model reported most on these cards is Qwen3.8 27B, mostly at IQ3_XXS, which leaves little room for context or runs partly from system RAM.

The calculator checks your exact card, including longer contexts and offloading.

Models that fit in 12GB

Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.

ModelBest quantNeeds
DeepSeek-Coder-V2-Lite 16B · 2.4B activeQ3_K_M9 GB
Qwen3 14BQ3_K_M9.2 GB
Qwen2.5 14BQ3_K_M9.2 GB
Qwen3.5 13BQ4_K_M9.4 GB
Gemma 4 12BQ4_K_M8.9 GB
Qwen3.5 9BQ6_K9.4 GB
DeepSeek V4 Flash 9BQ6_K9.5 GB
GLM-4 9BQ6_K9.4 GB
Gemma 2 9BQ6_K10 GB
Mimo 2.6 9BQ6_K9.3 GB
Ornith1.5 9BQ6_K9.3 GB
Gemma 4 8BQ8_010.1 GB

And 22 smaller models. The calculator lists them all for your card.

What people run on 12GB

Model sizes reported on setups with more than 8 GB and up to 12 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.

#ModelReports
1Qwen3.827B
Alibaba · mostly IQ3_XXS
Typical 20 t/s (4 runs) · with speculation 20 t/s (4)
19
2Qwen3.8 Flash-Next125B · 6B active
Alibaba · mostly IQ3_XXS
No plain runs
14
3Qwen3.635B · 3B active
Alibaba · mostly IQ2_XXS
No plain runs
8
4Qwen3.627B
Alibaba · mostly Q2-XS
Typical 24 t/s (1 run)
3
5Bonsai 2 Ternary27B
PrismML · mostly 2-bit ternary
Typical 60 t/s (1 run)
2
6Gemma 4 E4B8B
Google DeepMind · mostly Q4_K_M
Typical 50 t/s (2 runs)
2
7Gemma 412B
Google DeepMind · mostly IQ4_NL
Typical 35 t/s (1 run)
2
8DeepSeek V4 Flash284B · 13B active
DeepSeek · mostly IQ2_M
No plain runs
2
9Qwen38B
Alibaba · mostly Q4_K_M
Typical 61 t/s (1 run)
1
10Gemma 4 E2B5.1B
Google DeepMind · mostly Q4_K_M
Typical 60 t/s (1 run)
1
11Bonsai 227B
PrismML · mostly PTQ1_0
Typical 52 t/s (1 run)
1
12Qwen2.57B
Alibaba · mostly Q4_K_M
Typical 49 t/s (1 run)
1
13GLM-4.7 Flash30B
Zhipu AI
Typical 43 t/s (1 run)
1
14Muse Glimmer30B
Meta · mostly EXL3-SC 3.00bpw H4
Typical 30 t/s (1 run)
1
15Agents-A1 Millie35B · 3B active
InternScience · mostly ternary
Typical 22 t/s (1 run)
1
16GLM-49B
mostly Q4_0
Typical 21 t/s (1 run)
1
17KAT-Coder V2.5-Dev35B · 3B active
Kwaipilot
Typical 14 t/s (1 run)
1
18Qwen34B
Alibaba · mostly Q4_K_M
Typical 13 t/s (1 run)
1
19DeepSeek V4 Flash284B
DeepSeek · mostly W8A8
No plain runs
1
20DeepSeek V4 Flash REAP-150B150B
DeepSeek · mostly FP4
No plain runs
1
21Qwen3.59B
Alibaba · mostly Q4
No plain runs
1
22Qwen3.513B
Alibaba
No plain runs
1
23Qwen3.6 Uncensored-HauhauCS-Aggressive35B · 3B active
Alibaba · mostly Q3_K_P
No plain runs
1
24Qwen3.8 Bonsai-Llama-Jev27B
Alibaba · mostly Q2_64
No plain runs
1
25Qwen3.8 L0xRE27B
Alibaba · mostly IQ3
No plain runs · with speculation 40 t/s (1)
1

Frequently asked

What do people run on 12GB of VRAM?

On setups with more than 8 GB and up to 12 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 35B · 3B active. The table lists each with the quant most people used and the typical speed.

Does a Mac with 12GB count as 12GB of VRAM?

No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 12GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.