llamaperf

Best local LLMs for 24GB VRAM

How far 24GB of graphics memory goes for local models on cards like NVIDIA RTX 3090, NVIDIA RTX 4090, AMD RX 7900 XTX or NVIDIA Tesla P40 24GB: which models fit, and what people actually run on it.

Is 24GB of VRAM enough?

24GB holds a dense model of up to about 28B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Gemma 4 31B, which needs about 19.8 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.

The calculator checks your exact card, including longer contexts and offloading.

Models that fit in 24GB

Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.

And 44 smaller models. The calculator lists them all for your card.

What people run on 24GB

Model sizes reported on setups with more than 16 GB and up to 24 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.

#ModelReports
1Qwen3.827B
Alibaba · mostly Q4_K_M
Typical 35 t/s (16 runs) · with speculation 98 t/s (31)
101
2Qwen3.8 Flash-Next125B · 6B active
Alibaba · mostly UD-Q4_K_XL
No plain runs
27
3Qwen3.627B
Alibaba · mostly Q4_K_M
Typical 42 t/s (3 runs) · with speculation 50 t/s (6)
19
4DeepSeek V4 Flash284B · 13B active
DeepSeek · mostly 2-bit dynamic
No plain runs
15
5Qwen3.635B · 3B active
Alibaba · mostly IQ4
Typical 85 t/s (6 runs) · with speculation 52 t/s (1)
14
6Gemma 426B · 4B active
Google DeepMind · mostly Q4_0
Typical 140 t/s (4 runs)
6
7Muse Glimmer30B
Meta · mostly IQ4_XS
Typical 30 t/s (1 run) · with speculation 89 t/s (2)
4
8Gemma 431B
Google DeepMind · mostly Q4_K_S
Typical 33 t/s (2 runs)
3
9Qwen3.8 Swift-1.527B
Alibaba · mostly IQ4_XS
Typical 67 t/s (1 run) · with speculation 104 t/s (1)
2
10Qwen2.57B
Alibaba · mostly GPTQ-Int4
Typical 66 t/s (2 runs)
2
11Gemma 4 E4B8B
Google DeepMind · mostly FP16
Typical 58 t/s (1 run)
2
12Nemotron 3.5 Lightning30B · 3B active
NVIDIA · mostly Q4_0
No plain runs
2
13Qwen3.8 Swift-1.5-Uncensored27B
Alibaba · mostly IQ4_XS
No plain runs · with speculation 68 t/s (1)
2
14Qwen2.5 Coder-7B7B
Alibaba · mostly Q4_K_M
Typical 159 t/s (1 run)
1
15DeepSeek-Coder-V2-Lite16B · 2.4B active
mostly 4bit
Typical 126 t/s (1 run)
1
16Qwen330B · 3B active
Alibaba · mostly IQ3_XXS
Typical 74 t/s (1 run)
1
17Qwen3.8 Escha-W227B
Alibaba · mostly escha 2-bit
Typical 67 t/s (1 run)
1
18Qwen3.59B
Alibaba · mostly BF16
Typical 41 t/s (1 run)
1
19Qwen2.5 Coder-32B32B
Alibaba · mostly Q4_K_M
Typical 28 t/s (1 run)
1
20Qwen38B
Alibaba · mostly Q4_K_M
Typical 25 t/s (1 run)
1
21Qwen314B
Alibaba · mostly 4-bit
Typical 12 t/s (1 run)
1
22Qwen3.535B · 3B active
Alibaba · mostly Q4_K_XL
Typical 8.3 t/s (1 run)
1
23Qwen3 Heretic14B
Alibaba · mostly INT4_SYM
Typical 3.6 t/s (1 run)
1
24Bonsai 2 Ternary27B
PrismML · mostly Q4/Q5
No plain runs · with speculation 188 t/s (1)
1
25DeepSeek V4 Flash Vision-Exp285B · 13B active
DeepSeek · mostly FP4
No plain runs
1
26DeepSeek V4.1 Flash552B · 16B active
mostly 4-bit
No plain runs
1
27GLM-4.5-Air106B · 12B active
No plain runs
1
28GLM-5.2744B · 40B active
Zhipu AI · mostly UD-IQ2_M
No plain runs
1
29GLM-5.3 Flash320B · 18B active
Zhipu AI · mostly UD-Q4_K_XL
No plain runs
1
30Gemma 4 Uncensored-HauhauCS-Balanced26B · 4B active
Google DeepMind · mostly Q4_K_M
No plain runs
1
31Ornith1.5 BigBang35B · 3B active
mostly Q4_K_M
No plain runs · with speculation 169 t/s (1)
1
32Qwen2.532B
Alibaba · mostly AWQ
No plain runs
1
33Qwen3235B · 22B active
Alibaba · mostly Q4_K_M
No plain runs
1
34Qwen3-Next80B · 3B active
Alibaba · mostly Q4_K_M
No plain runs
1
35Qwen3.5 Mica v0.14B
Alibaba · mostly Q4_K_M
No plain runs
1
36Qwen3.54B
Alibaba
No plain runs
1
37Qwen3.6 Heretic-v227B
Alibaba · mostly Q4_K_M
No plain runs · with speculation 80 t/s (1)
1
38Qwen3.8 Swift-1.5-HyperQwen27B
Alibaba · mostly W4A16
No plain runs
1
39Qwen3.8 Swift27B
Alibaba · mostly IQ4_XS
No plain runs
1
40Qwen3.8 CODER27B
Alibaba · mostly IQ4_XS
No plain runs
1
41Qwen3.8 Huihui-Abliterated27B
Alibaba · mostly Q4_K_S
No plain runs
1
42Qwen3.8 Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO27B
Alibaba · mostly Q8_0
No plain runs
1
43Qwen3.8 heretic-ara27B
Alibaba · mostly Q5_K_M
No plain runs
1

Frequently asked

What do people run on 24GB of VRAM?

On setups with more than 16 GB and up to 24 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 27B. The table lists each with the quant most people used and the typical speed.

Does a Mac with 24GB count as 24GB of VRAM?

No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 24GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.