llamaperf

Best local LLMs for 32GB VRAM

How far 32GB of graphics memory goes for local models on cards like NVIDIA RTX 5090, AMD Radeon AI PRO R9700 32GB, NVIDIA V100 32GB or Intel Arc Pro B70: which models fit, and what people actually run on it.

Is 32GB of VRAM enough?

32GB holds a dense model of up to about 39B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.6 35B · 3B active, which needs about 26 GB at Q5_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.

The calculator checks your exact card, including longer contexts and offloading.

Models that fit in 32GB

Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.

And 44 smaller models. The calculator lists them all for your card.

What people run on 32GB

Model sizes reported on setups with more than 24 GB and up to 32 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.

#ModelReports
1Qwen3.827B
Alibaba · mostly NVFP4
Typical 31 t/s (18 runs) · with speculation 120 t/s (29)
97
2Qwen3.8 Flash-Next125B · 6B active
Alibaba · mostly Q4_K_M
Typical 59 t/s (2 runs)
34
3Qwen3.635B · 3B active
Alibaba · mostly Q4_K_M
Typical 63 t/s (12 runs)
17
4Qwen3.627B
Alibaba · mostly NVFP4
Typical 39 t/s (2 runs) · with speculation 77 t/s (2)
9
5DeepSeek V4 Flash284B · 13B active
DeepSeek · mostly Q8_K_XL
No plain runs
9
6Gemma 426B · 4B active
Google DeepMind · mostly Q4_K_M
Typical 148 t/s (2 runs) · with speculation 578 t/s (1)
7
7Ornith1.535B · 3B active
mostly MXFP4
Typical 142 t/s (2 runs) · with speculation 245 t/s (1)
5
8Muse Glimmer30B
Meta · mostly Q5_K_XL
No plain runs · with speculation 243 t/s (2)
4
9Qwen3.8 Swift27B
Alibaba · mostly Q6
No plain runs
3
10Qwen3.8 Huihui-Abliterated27B
Alibaba · mostly NVFP4
No plain runs · with speculation 201 t/s (3)
3
11Qwen3.8 Flash-Next Uncensored125B · 6B active
Alibaba · mostly NVFP4
No plain runs
3
12Qwen3.535B · 3B active
Alibaba · mostly Q4_K_XL
Typical 172 t/s (2 runs)
2
13Qwen38B
Alibaba · mostly BF16
Typical 78 t/s (1 run) · with speculation 159 t/s (1)
2
14Qwen3.8 Uncensored27B
Alibaba · mostly NVFP4
Typical 77 t/s (1 run) · with speculation 175 t/s (1)
2
15Qwen3.8 Swift-1.527B
Alibaba · mostly MXFP4
Typical 38 t/s (1 run) · with speculation 161 t/s (1)
2
16DeepSeek V4.1 Flash552B · 16B active
mostly MXFP4
No plain runs
2
17Nemotron 3.5 Lightning30B · 3B active
NVIDIA · mostly NVFP4
No plain runs
2
18Qwen2.57B
Alibaba · mostly AWQ
No plain runs
2
19Qwen3.54B
Alibaba · mostly Q8_0
Typical 107 t/s (1 run)
1
20Bonsai 227B
PrismML · mostly PQ2_0
Typical 101 t/s (1 run)
1
21Qwen3-Coder30B · 3B active
mostly Q4_K_M
Typical 66 t/s (1 run)
1
22Qwen3 Thinking-250730B · 3B active
Alibaba · mostly Q4_0
Typical 63 t/s (1 run)
1
23Qwen330B · 3B active
Alibaba · mostly 4-bit
Typical 57 t/s (1 run)
1
24Bonsai 2 Ternary27B
PrismML · mostly PTQ1_0
Typical 37 t/s (1 run)
1
25GLM-5.3 Flash320B · 18B active
Zhipu AI · mostly FP8-E4M3
No plain runs
1
26Gemma 412B
Google DeepMind
No plain runs
1
27Kimi K2.51000B · 32B active
Moonshot AI · mostly IQ3_M
No plain runs
1
28Kimi K2.61000B · 32B active
Moonshot AI
No plain runs
1
29Qwen3.59B
Alibaba
No plain runs
1
30Qwen3.6 Uncensored-HauhauCS-Aggressive35B · 3B active
Alibaba · mostly Q4_K_P
No plain runs
1
31Qwen3.8 Swift Uncensored27B
Alibaba · mostly Q6_K
No plain runs
1
32Qwen3.8 RVN Heretic (ARA abliterated)27B
Alibaba · mostly Q6_K
No plain runs
1
33Qwen3.8 Swift-Genesis27B
Alibaba
No plain runs
1
34Qwen3.8 QUASAR27B
Alibaba · mostly NVFP4
No plain runs · with speculation 287 t/s (1)
1
35Qwen3.8 Uncensored-HauhauCS-Aggressive27B
Alibaba · mostly Q5_K_P
No plain runs · with speculation 14 t/s (1)
1
36Qwen3.8 Swift-1.5-125B-A6B125B · 6B active
Alibaba · mostly IQ2_XS
No plain runs
1
37Qwen3.82400B · 95B active
Alibaba · mostly UD-Q1_0
No plain runs
1
38Qwopus3.6 Coder-Compat27B
mostly Q4_K_M
No plain runs · with speculation 40 t/s (1)
1

Frequently asked

What do people run on 32GB of VRAM?

On setups with more than 24 GB and up to 32 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 35B · 3B active. The table lists each with the quant most people used and the typical speed.

Does a Mac with 32GB count as 32GB of VRAM?

No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 32GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.