Best local LLMs for 32GB VRAM
How far 32GB of graphics memory goes for local models on cards like NVIDIA RTX 5090, AMD Radeon AI PRO R9700 32GB, NVIDIA V100 32GB or Intel Arc Pro B70: which models fit, and what people actually run on it.
Is 32GB of VRAM enough?
32GB holds a dense model of up to about 39B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.6 35B · 3B active, which needs about 26 GB at Q5_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 32GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| Qwen3.6 35B · 3B active | Q5_K_M | 26 GB |
| Ornith1.5 35B · 3B active | Q5_K_M | 26 GB |
| Qwen3.5 35B · 3B active | Q5_K_M | 26 GB |
| Agents-A1 35B · 3B active | Q5_K_M | 26 GB |
| Nex-N2.5-mini 35B · 3B active | Q5_K_M | 26 GB |
| KAT-Coder V2.5-Dev 35B · 3B active | Q5_K_M | 26 GB |
| Qwen2.5 32B | Q5_K_M | 25.9 GB |
| Qwen3 32B | Q5_K_M | 25.9 GB |
| Gemma 4 31B | Q5_K_M | 23.7 GB |
| Muse 30B | Q6_K | 26.3 GB |
| Qwen3 30B · 3B active | Q5_K_M | 24.5 GB |
| Nemotron 3.5 Lightning 30B · 3B active | Q6_K | 26.2 GB |
And 44 smaller models. The calculator lists them all for your card.
What people run on 32GB
Model sizes reported on setups with more than 24 GB and up to 32 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.827B Alibaba · mostly NVFP4 Typical 31 t/s (18 runs) · with speculation 120 t/s (29) | 97 | 31 18 runs, 7 devices | 120 29 runs |
| 2 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly Q4_K_M Typical 59 t/s (2 runs) | 34 | 59 2 runs | none |
| 3 | Qwen3.635B · 3B active Alibaba · mostly Q4_K_M Typical 63 t/s (12 runs) | 17 | 63 12 runs, 6 devices | none |
| 4 | Qwen3.627B Alibaba · mostly NVFP4 Typical 39 t/s (2 runs) · with speculation 77 t/s (2) | 9 | 39 2 runs, 2 devices | 77 2 runs |
| 5 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly Q8_K_XL No plain runs | 9 | no plain runs | none |
| 6 | Gemma 426B · 4B active Google DeepMind · mostly Q4_K_M Typical 148 t/s (2 runs) · with speculation 578 t/s (1) | 7 | 148 2 runs, 2 devices | 578 1 run |
| 7 | Ornith1.535B · 3B active mostly MXFP4 Typical 142 t/s (2 runs) · with speculation 245 t/s (1) | 5 | 142 2 runs, 2 devices | 245 1 run |
| 8 | Muse Glimmer30B Meta · mostly Q5_K_XL No plain runs · with speculation 243 t/s (2) | 4 | no plain runs | 243 2 runs |
| 9 | Qwen3.8 Swift27B Alibaba · mostly Q6 No plain runs | 3 | no plain runs | none |
| 10 | Qwen3.8 Huihui-Abliterated27B Alibaba · mostly NVFP4 No plain runs · with speculation 201 t/s (3) | 3 | no plain runs | 201 3 runs |
| 11 | Qwen3.8 Flash-Next Uncensored125B · 6B active Alibaba · mostly NVFP4 No plain runs | 3 | no plain runs | none |
| 12 | Qwen3.535B · 3B active Alibaba · mostly Q4_K_XL Typical 172 t/s (2 runs) | 2 | 172 2 runs, 2 devices | none |
| 13 | Qwen38B Alibaba · mostly BF16 Typical 78 t/s (1 run) · with speculation 159 t/s (1) | 2 | 78 1 run | 159 1 run |
| 14 | Qwen3.8 Uncensored27B Alibaba · mostly NVFP4 Typical 77 t/s (1 run) · with speculation 175 t/s (1) | 2 | 77 1 run | 175 1 run |
| 15 | Qwen3.8 Swift-1.527B Alibaba · mostly MXFP4 Typical 38 t/s (1 run) · with speculation 161 t/s (1) | 2 | 38 1 run | 161 1 run |
| 16 | DeepSeek V4.1 Flash552B · 16B active mostly MXFP4 No plain runs | 2 | no plain runs | none |
| 17 | Nemotron 3.5 Lightning30B · 3B active NVIDIA · mostly NVFP4 No plain runs | 2 | no plain runs | none |
| 18 | Qwen2.57B Alibaba · mostly AWQ No plain runs | 2 | no plain runs | none |
| 19 | Qwen3.54B Alibaba · mostly Q8_0 Typical 107 t/s (1 run) | 1 | 107 1 run | none |
| 20 | Bonsai 227B PrismML · mostly PQ2_0 Typical 101 t/s (1 run) | 1 | 101 1 run | none |
| 21 | Qwen3-Coder30B · 3B active mostly Q4_K_M Typical 66 t/s (1 run) | 1 | 66 1 run | none |
| 22 | Qwen3 Thinking-250730B · 3B active Alibaba · mostly Q4_0 Typical 63 t/s (1 run) | 1 | 63 1 run | none |
| 23 | Qwen330B · 3B active Alibaba · mostly 4-bit Typical 57 t/s (1 run) | 1 | 57 1 run | none |
| 24 | Bonsai 2 Ternary27B PrismML · mostly PTQ1_0 Typical 37 t/s (1 run) | 1 | 37 1 run | none |
| 25 | GLM-5.3 Flash320B · 18B active Zhipu AI · mostly FP8-E4M3 No plain runs | 1 | no plain runs | none |
| 26 | Gemma 412B Google DeepMind No plain runs | 1 | no plain runs | none |
| 27 | Kimi K2.51000B · 32B active Moonshot AI · mostly IQ3_M No plain runs | 1 | no plain runs | none |
| 28 | Kimi K2.61000B · 32B active Moonshot AI No plain runs | 1 | no plain runs | none |
| 29 | Qwen3.59B Alibaba No plain runs | 1 | no plain runs | none |
| 30 | Qwen3.6 Uncensored-HauhauCS-Aggressive35B · 3B active Alibaba · mostly Q4_K_P No plain runs | 1 | no plain runs | none |
| 31 | Qwen3.8 Swift Uncensored27B Alibaba · mostly Q6_K No plain runs | 1 | no plain runs | none |
| 32 | Qwen3.8 RVN Heretic (ARA abliterated)27B Alibaba · mostly Q6_K No plain runs | 1 | no plain runs | none |
| 33 | Qwen3.8 Swift-Genesis27B Alibaba No plain runs | 1 | no plain runs | none |
| 34 | Qwen3.8 QUASAR27B Alibaba · mostly NVFP4 No plain runs · with speculation 287 t/s (1) | 1 | no plain runs | 287 1 run |
| 35 | Qwen3.8 Uncensored-HauhauCS-Aggressive27B Alibaba · mostly Q5_K_P No plain runs · with speculation 14 t/s (1) | 1 | no plain runs | 14 1 run |
| 36 | Qwen3.8 Swift-1.5-125B-A6B125B · 6B active Alibaba · mostly IQ2_XS No plain runs | 1 | no plain runs | none |
| 37 | Qwen3.82400B · 95B active Alibaba · mostly UD-Q1_0 No plain runs | 1 | no plain runs | none |
| 38 | Qwopus3.6 Coder-Compat27B mostly Q4_K_M No plain runs · with speculation 40 t/s (1) | 1 | no plain runs | 40 1 run |
Frequently asked
What do people run on 32GB of VRAM?
On setups with more than 24 GB and up to 32 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 35B · 3B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 32GB count as 32GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 32GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.