Best local LLMs for 8GB VRAM
How far 8GB of graphics memory goes for local models on cards like NVIDIA RTX 4060 Ti 8GB, NVIDIA RTX 5060 8GB, NVIDIA RTX 5060 Laptop 8GB or AMD RX 6600 XT: which models fit, and what people actually run on it.
Is 8GB of VRAM enough?
8GB holds a dense model of up to about 6B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Gemma 4 8B, which needs about 6.6 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top. People do push further: the model reported most on these cards is Qwen3.6 35B · 3B active, mostly at Q4_K_XL, which leaves little room for context or runs partly from system RAM.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 8GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| Qwen3.5 9B | Q3_K_M | 6 GB |
| DeepSeek V4 Flash 9B | Q3_K_M | 6.1 GB |
| GLM-4 9B | Q3_K_M | 6 GB |
| Gemma 2 9B | Q3_K_M | 6.6 GB |
| Mimo 2.6 9B | Q3_K_M | 5.9 GB |
| Ornith1.5 9B | Q3_K_M | 6 GB |
| Gemma 4 8B | Q4_K_M | 6.6 GB |
| Qwen3 8B | Q3_K_M | 6.5 GB |
| Llama 3.1 8B | Q3_K_M | 6.5 GB |
| Ling-3.0 7.9B · 1.3B active | Q4_K_M | 6.2 GB |
| Qwen2.5 7B | Q3_K_M | 6 GB |
| Gemma 4 5.1B | Q6_K | 6.2 GB |
And 17 smaller models. The calculator lists them all for your card.
What people run on 8GB
Model sizes reported on setups with up to 8 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.635B · 3B active Alibaba · mostly Q4_K_XL No plain runs | 10 | no plain runs | none |
| 2 | Ornith1.535B · 3B active mostly Q4_K_M No plain runs | 4 | no plain runs | none |
| 3 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly 1 bit No plain runs | 4 | no plain runs | none |
| 4 | Bonsai 227B PrismML · mostly PTQ1_0 Typical 7.6 t/s (3 runs) | 3 | 7.6 3 runs, 3 devices | none |
| 5 | Qwen2.57B Alibaba Typical 32 t/s (2 runs) | 2 | 32 2 runs | none |
| 6 | Qwen3.827B Alibaba · mostly UD-IQ2_XXS Typical 30 t/s (1 run) | 2 | 30 1 run | none |
| 7 | Gemma 426B · 4B active Google DeepMind · mostly QAT No plain runs | 2 | no plain runs | none |
| 8 | LFM2.51.2B Liquid AI Typical 129 t/s (1 run) | 1 | 129 1 run | none |
| 9 | Qwen2.5 Coder-1.5B1.5B Alibaba Typical 99 t/s (1 run) | 1 | 99 1 run | none |
| 10 | Qwen30.6B Alibaba Typical 76 t/s (1 run) | 1 | 76 1 run | none |
| 11 | Ling-3.0 Tiny7.9B · 1.3B active Ant Group · mostly Q4_K_M Typical 54 t/s (1 run) | 1 | 54 1 run | none |
| 12 | LFM2.52.6B Liquid AI · mostly QAD-Q4_0 Typical 22 t/s (1 run) | 1 | 22 1 run | none |
| 13 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly MXFP4 Typical 3.2 t/s (1 run) | 1 | 3.2 1 run | none |
| 14 | Bonsai27B PrismML · mostly 1-bit No plain runs | 1 | no plain runs | none |
| 15 | DeepSeek V4.1 Flash552B · 16B active mostly FP8 No plain runs | 1 | no plain runs | none |
| 16 | Gemma 4 E2B5.1B Google DeepMind · mostly Q4 No plain runs | 1 | no plain runs | none |
| 17 | Gemma 4 E4B8B Google DeepMind No plain runs | 1 | no plain runs | none |
| 18 | Gemma 412B Google DeepMind · mostly Q6_K No plain runs | 1 | no plain runs | none |
| 19 | Gemma 4 Uncensored26B · 4B active Google DeepMind No plain runs | 1 | no plain runs | none |
| 20 | Gemma 4 Ultra-Uncensored-Heretic26B · 4B active Google DeepMind · mostly IQ4_XS No plain runs | 1 | no plain runs | none |
| 21 | Qwen2.51.5B Alibaba · mostly Q2_K No plain runs | 1 | no plain runs | none |
| 22 | Qwen3.59B Alibaba No plain runs | 1 | no plain runs | none |
| 23 | Qwen3.535B · 3B active Alibaba · mostly Q2 No plain runs | 1 | no plain runs | none |
| 24 | Qwen3.627B Alibaba · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 25 | Qwen3.8 Swift27B Alibaba · mostly W4A16 No plain runs | 1 | no plain runs | none |
Frequently asked
What do people run on 8GB of VRAM?
On setups with up to 8 GB of memory, the model people report most is Qwen3.6 35B · 3B active, followed by Ornith1.5 35B · 3B active and Qwen3.8 125B · 6B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 8GB count as 8GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 8GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.