Best local LLMs for 16GB VRAM
How far 16GB of graphics memory goes for local models on cards like NVIDIA RTX 5060 Ti 16GB, NVIDIA RTX 5070 Ti, NVIDIA RTX 5080 or NVIDIA V100 16GB: which models fit, and what people actually run on it.
Is 16GB of VRAM enough?
16GB holds a dense model of up to about 16B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is DeepSeek-Coder-V2-Lite 16B · 2.4B active, which needs about 13 GB at Q5_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top. People do push further: the model reported most on these cards is Qwen3.8 27B, mostly at IQ4_XS, which leaves little room for context or runs partly from system RAM.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 16GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| DeepSeek-Coder-V2-Lite 16B · 2.4B active | Q5_K_M | 13 GB |
| Qwen3 14B | Q5_K_M | 12.7 GB |
| Qwen2.5 14B | Q5_K_M | 12.7 GB |
| Qwen3.5 13B | Q6_K | 12.7 GB |
| Gemma 4 12B | Q6_K | 11.9 GB |
| Qwen3.5 9B | Q8_0 | 11.1 GB |
| DeepSeek V4 Flash 9B | Q8_0 | 11.2 GB |
| GLM-4 9B | Q8_0 | 11.1 GB |
| Gemma 2 9B | Q8_0 | 11.7 GB |
| Mimo 2.6 9B | Q8_0 | 11 GB |
| Ornith1.5 9B | Q8_0 | 11 GB |
| Gemma 4 8B | Q8_0 | 10.1 GB |
And 22 smaller models. The calculator lists them all for your card.
What people run on 16GB
Model sizes reported on setups with more than 12 GB and up to 16 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.827B Alibaba · mostly IQ4_XS Typical 25 t/s (12 runs) · with speculation 56 t/s (19) | 69 | 25 12 runs, 9 devices | 56 19 runs |
| 2 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly UD-Q4_K_XL No plain runs | 21 | no plain runs | none |
| 3 | Qwen3.635B · 3B active Alibaba · mostly Q4_K_XL No plain runs · with speculation 110 t/s (1) | 15 | no plain runs | 110 1 run |
| 4 | Qwen3.627B Alibaba · mostly IQ4_XS Typical 22 t/s (1 run) | 4 | 22 1 run | none |
| 5 | Gemma 426B · 4B active Google DeepMind · mostly IQ4_XS Typical 63 t/s (2 runs) | 3 | 63 2 runs, 2 devices | none |
| 6 | Qwen3.8 Swift-1.527B Alibaba · mostly IQ3_S Typical 42 t/s (2 runs) · with speculation 44 t/s (1) | 3 | 42 2 runs, 2 devices | 44 1 run |
| 7 | Qwen330B · 3B active Alibaba · mostly Q4_K_M No plain runs | 3 | no plain runs | none |
| 8 | Qwen2.50.5B Alibaba · mostly Q4 Typical 96 t/s (1 run) | 2 | 96 1 run | none |
| 9 | Gemma 4 E4B8B Google DeepMind · mostly 4bit Typical 41 t/s (2 runs) | 2 | 41 2 runs, 2 devices | none |
| 10 | Qwen3.59B Alibaba · mostly Q6_K Typical 33 t/s (2 runs) | 2 | 33 2 runs, 2 devices | none |
| 11 | Muse Glimmer30B Meta · mostly Q4_K_XL Typical 18 t/s (1 run) · with speculation 20 t/s (1) | 2 | 18 1 run | 20 1 run |
| 12 | Qwen3.535B · 3B active Alibaba · mostly UD-IQ3_XXS Typical 17 t/s (1 run) | 2 | 17 1 run | none |
| 13 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly UD-Q2_K_XL No plain runs | 2 | no plain runs | none |
| 14 | Gemma 426B Google DeepMind · mostly EXL3 like No plain runs · with speculation 182 t/s (1) | 2 | no plain runs | 182 1 run |
| 15 | Qwen3.8 Uncensored27B Alibaba · mostly IQ3_XXS No plain runs · with speculation 35 t/s (1) | 2 | no plain runs | 35 1 run |
| 16 | Bonsai27B PrismML · mostly Q1_0 Typical 73 t/s (1 run) | 1 | 73 1 run | none |
| 17 | Qwen3.8 Escha-W227B Alibaba · mostly 2.469 bpw Typical 59 t/s (1 run) | 1 | 59 1 run | none |
| 18 | Gemma 412B Google DeepMind · mostly Q5_K_XL Typical 50 t/s (1 run) | 1 | 50 1 run | none |
| 19 | Mimo 2.6 Distill-Qwen9B Xiaomi · mostly Q8_0 Typical 42 t/s (1 run) | 1 | 42 1 run | none |
| 20 | Qwen3.6 APEX-I-Quality35B · 3B active Alibaba Typical 37 t/s (1 run) | 1 | 37 1 run | none |
| 21 | Qwen2.5 Coder14B Alibaba · mostly Q4_K_M Typical 34 t/s (1 run) | 1 | 34 1 run | none |
| 22 | Qwen31.7B Alibaba · mostly Q8_0 Typical 33 t/s (1 run) | 1 | 33 1 run | none |
| 23 | Qwen3.6 Coder27B · 3B active Alibaba · mostly Q4_K_M Typical 14 t/s (1 run) | 1 | 14 1 run | none |
| 24 | Llama 3.18B Meta · mostly Q4_K_M Typical 1.5 t/s (1 run) | 1 | 1.5 1 run | none |
| 25 | Agents-A1 Uncensored-MTP-APEX35B · 3B active InternScience · mostly APEX Compact No plain runs · with speculation 97 t/s (1) | 1 | no plain runs | 97 1 run |
| 26 | Ornith1.59B mostly Q6_K No plain runs | 1 | no plain runs | none |
| 27 | Qwen3.5 DeepSeek-V4-Flash9B Alibaba · mostly Q6_K No plain runs · with speculation 49 t/s (1) | 1 | no plain runs | 49 1 run |
| 28 | Qwen3.8 abliterated27B Alibaba · mostly IQ3_S No plain runs · with speculation 105 t/s (1) | 1 | no plain runs | 105 1 run |
| 29 | Qwen3.8 TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX27B Alibaba No plain runs · with speculation 30 t/s (1) | 1 | no plain runs | 30 1 run |
| 30 | Qwen3.8 Huihui-Abliterated27B Alibaba · mostly IQ3_XXS No plain runs | 1 | no plain runs | none |
| 31 | Qwen3.8 Swift-1.5 Flash-Next125B · 6B active Alibaba · mostly IQ2_XS No plain runs | 1 | no plain runs | none |
Frequently asked
What do people run on 16GB of VRAM?
On setups with more than 12 GB and up to 16 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 35B · 3B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 16GB count as 16GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 16GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.