Best local LLMs for 12GB VRAM
How far 12GB of graphics memory goes for local models on cards like NVIDIA RTX 3060 12GB, NVIDIA RTX 4070, NVIDIA RTX 5070 or NVIDIA RTX 5070 Ti Laptop 12GB: which models fit, and what people actually run on it.
Is 12GB of VRAM enough?
12GB holds a dense model of up to about 12B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.5 13B, which needs about 9.4 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top. People do push further: the model reported most on these cards is Qwen3.8 27B, mostly at IQ3_XXS, which leaves little room for context or runs partly from system RAM.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 12GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| DeepSeek-Coder-V2-Lite 16B · 2.4B active | Q3_K_M | 9 GB |
| Qwen3 14B | Q3_K_M | 9.2 GB |
| Qwen2.5 14B | Q3_K_M | 9.2 GB |
| Qwen3.5 13B | Q4_K_M | 9.4 GB |
| Gemma 4 12B | Q4_K_M | 8.9 GB |
| Qwen3.5 9B | Q6_K | 9.4 GB |
| DeepSeek V4 Flash 9B | Q6_K | 9.5 GB |
| GLM-4 9B | Q6_K | 9.4 GB |
| Gemma 2 9B | Q6_K | 10 GB |
| Mimo 2.6 9B | Q6_K | 9.3 GB |
| Ornith1.5 9B | Q6_K | 9.3 GB |
| Gemma 4 8B | Q8_0 | 10.1 GB |
And 22 smaller models. The calculator lists them all for your card.
What people run on 12GB
Model sizes reported on setups with more than 8 GB and up to 12 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.827B Alibaba · mostly IQ3_XXS Typical 20 t/s (4 runs) · with speculation 20 t/s (4) | 19 | 20 4 runs, 3 devices | 20 4 runs |
| 2 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly IQ3_XXS No plain runs | 14 | no plain runs | none |
| 3 | Qwen3.635B · 3B active Alibaba · mostly IQ2_XXS No plain runs | 8 | no plain runs | none |
| 4 | Qwen3.627B Alibaba · mostly Q2-XS Typical 24 t/s (1 run) | 3 | 24 1 run | none |
| 5 | Bonsai 2 Ternary27B PrismML · mostly 2-bit ternary Typical 60 t/s (1 run) | 2 | 60 1 run | none |
| 6 | Gemma 4 E4B8B Google DeepMind · mostly Q4_K_M Typical 50 t/s (2 runs) | 2 | 50 2 runs, 2 devices | none |
| 7 | Gemma 412B Google DeepMind · mostly IQ4_NL Typical 35 t/s (1 run) | 2 | 35 1 run | none |
| 8 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly IQ2_M No plain runs | 2 | no plain runs | none |
| 9 | Qwen38B Alibaba · mostly Q4_K_M Typical 61 t/s (1 run) | 1 | 61 1 run | none |
| 10 | Gemma 4 E2B5.1B Google DeepMind · mostly Q4_K_M Typical 60 t/s (1 run) | 1 | 60 1 run | none |
| 11 | Bonsai 227B PrismML · mostly PTQ1_0 Typical 52 t/s (1 run) | 1 | 52 1 run | none |
| 12 | Qwen2.57B Alibaba · mostly Q4_K_M Typical 49 t/s (1 run) | 1 | 49 1 run | none |
| 13 | GLM-4.7 Flash30B Zhipu AI Typical 43 t/s (1 run) | 1 | 43 1 run | none |
| 14 | Muse Glimmer30B Meta · mostly EXL3-SC 3.00bpw H4 Typical 30 t/s (1 run) | 1 | 30 1 run | none |
| 15 | Agents-A1 Millie35B · 3B active InternScience · mostly ternary Typical 22 t/s (1 run) | 1 | 22 1 run | none |
| 16 | GLM-49B mostly Q4_0 Typical 21 t/s (1 run) | 1 | 21 1 run | none |
| 17 | KAT-Coder V2.5-Dev35B · 3B active Kwaipilot Typical 14 t/s (1 run) | 1 | 14 1 run | none |
| 18 | Qwen34B Alibaba · mostly Q4_K_M Typical 13 t/s (1 run) | 1 | 13 1 run | none |
| 19 | DeepSeek V4 Flash284B DeepSeek · mostly W8A8 No plain runs | 1 | no plain runs | none |
| 20 | DeepSeek V4 Flash REAP-150B150B DeepSeek · mostly FP4 No plain runs | 1 | no plain runs | none |
| 21 | Qwen3.59B Alibaba · mostly Q4 No plain runs | 1 | no plain runs | none |
| 22 | Qwen3.513B Alibaba No plain runs | 1 | no plain runs | none |
| 23 | Qwen3.6 Uncensored-HauhauCS-Aggressive35B · 3B active Alibaba · mostly Q3_K_P No plain runs | 1 | no plain runs | none |
| 24 | Qwen3.8 Bonsai-Llama-Jev27B Alibaba · mostly Q2_64 No plain runs | 1 | no plain runs | none |
| 25 | Qwen3.8 L0xRE27B Alibaba · mostly IQ3 No plain runs · with speculation 40 t/s (1) | 1 | no plain runs | 40 1 run |
Frequently asked
What do people run on 12GB of VRAM?
On setups with more than 8 GB and up to 12 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 35B · 3B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 12GB count as 12GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 12GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.