llamaperf

Best local LLMs for 16GB VRAM

How far 16GB of graphics memory goes for local models on cards like NVIDIA RTX 5060 Ti 16GB, NVIDIA RTX 5070 Ti, NVIDIA RTX 5080 or NVIDIA V100 16GB: which models fit, and what people actually run on it.

Is 16GB of VRAM enough?

16GB holds a dense model of up to about 16B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is DeepSeek-Coder-V2-Lite 16B · 2.4B active, which needs about 13 GB at Q5_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top. People do push further: the model reported most on these cards is Qwen3.8 27B, mostly at IQ4_XS, which leaves little room for context or runs partly from system RAM.

The calculator checks your exact card, including longer contexts and offloading.

Models that fit in 16GB

Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.

ModelBest quantNeeds
DeepSeek-Coder-V2-Lite 16B · 2.4B activeQ5_K_M13 GB
Qwen3 14BQ5_K_M12.7 GB
Qwen2.5 14BQ5_K_M12.7 GB
Qwen3.5 13BQ6_K12.7 GB
Gemma 4 12BQ6_K11.9 GB
Qwen3.5 9BQ8_011.1 GB
DeepSeek V4 Flash 9BQ8_011.2 GB
GLM-4 9BQ8_011.1 GB
Gemma 2 9BQ8_011.7 GB
Mimo 2.6 9BQ8_011 GB
Ornith1.5 9BQ8_011 GB
Gemma 4 8BQ8_010.1 GB

And 22 smaller models. The calculator lists them all for your card.

What people run on 16GB

Model sizes reported on setups with more than 12 GB and up to 16 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.

#ModelReports
1Qwen3.827B
Alibaba · mostly IQ4_XS
Typical 25 t/s (12 runs) · with speculation 56 t/s (19)
69
2Qwen3.8 Flash-Next125B · 6B active
Alibaba · mostly UD-Q4_K_XL
No plain runs
21
3Qwen3.635B · 3B active
Alibaba · mostly Q4_K_XL
No plain runs · with speculation 110 t/s (1)
15
4Qwen3.627B
Alibaba · mostly IQ4_XS
Typical 22 t/s (1 run)
4
5Gemma 426B · 4B active
Google DeepMind · mostly IQ4_XS
Typical 63 t/s (2 runs)
3
6Qwen3.8 Swift-1.527B
Alibaba · mostly IQ3_S
Typical 42 t/s (2 runs) · with speculation 44 t/s (1)
3
7Qwen330B · 3B active
Alibaba · mostly Q4_K_M
No plain runs
3
8Qwen2.50.5B
Alibaba · mostly Q4
Typical 96 t/s (1 run)
2
9Gemma 4 E4B8B
Google DeepMind · mostly 4bit
Typical 41 t/s (2 runs)
2
10Qwen3.59B
Alibaba · mostly Q6_K
Typical 33 t/s (2 runs)
2
11Muse Glimmer30B
Meta · mostly Q4_K_XL
Typical 18 t/s (1 run) · with speculation 20 t/s (1)
2
12Qwen3.535B · 3B active
Alibaba · mostly UD-IQ3_XXS
Typical 17 t/s (1 run)
2
13DeepSeek V4 Flash284B · 13B active
DeepSeek · mostly UD-Q2_K_XL
No plain runs
2
14Gemma 426B
Google DeepMind · mostly EXL3 like
No plain runs · with speculation 182 t/s (1)
2
15Qwen3.8 Uncensored27B
Alibaba · mostly IQ3_XXS
No plain runs · with speculation 35 t/s (1)
2
16Bonsai27B
PrismML · mostly Q1_0
Typical 73 t/s (1 run)
1
17Qwen3.8 Escha-W227B
Alibaba · mostly 2.469 bpw
Typical 59 t/s (1 run)
1
18Gemma 412B
Google DeepMind · mostly Q5_K_XL
Typical 50 t/s (1 run)
1
19Mimo 2.6 Distill-Qwen9B
Xiaomi · mostly Q8_0
Typical 42 t/s (1 run)
1
20Qwen3.6 APEX-I-Quality35B · 3B active
Alibaba
Typical 37 t/s (1 run)
1
21Qwen2.5 Coder14B
Alibaba · mostly Q4_K_M
Typical 34 t/s (1 run)
1
22Qwen31.7B
Alibaba · mostly Q8_0
Typical 33 t/s (1 run)
1
23Qwen3.6 Coder27B · 3B active
Alibaba · mostly Q4_K_M
Typical 14 t/s (1 run)
1
24Llama 3.18B
Meta · mostly Q4_K_M
Typical 1.5 t/s (1 run)
1
25Agents-A1 Uncensored-MTP-APEX35B · 3B active
InternScience · mostly APEX Compact
No plain runs · with speculation 97 t/s (1)
1
26Ornith1.59B
mostly Q6_K
No plain runs
1
27Qwen3.5 DeepSeek-V4-Flash9B
Alibaba · mostly Q6_K
No plain runs · with speculation 49 t/s (1)
1
28Qwen3.8 abliterated27B
Alibaba · mostly IQ3_S
No plain runs · with speculation 105 t/s (1)
1
29Qwen3.8 TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX27B
Alibaba
No plain runs · with speculation 30 t/s (1)
1
30Qwen3.8 Huihui-Abliterated27B
Alibaba · mostly IQ3_XXS
No plain runs
1
31Qwen3.8 Swift-1.5 Flash-Next125B · 6B active
Alibaba · mostly IQ2_XS
No plain runs
1

Frequently asked

What do people run on 16GB of VRAM?

On setups with more than 12 GB and up to 16 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 35B · 3B active. The table lists each with the quant most people used and the typical speed.

Does a Mac with 16GB count as 16GB of VRAM?

No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 16GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.