Best local LLMs for 24GB VRAM
How far 24GB of graphics memory goes for local models on cards like NVIDIA RTX 3090, NVIDIA RTX 4090, AMD RX 7900 XTX or NVIDIA Tesla P40 24GB: which models fit, and what people actually run on it.
Is 24GB of VRAM enough?
24GB holds a dense model of up to about 28B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Gemma 4 31B, which needs about 19.8 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 24GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| Qwen3.6 35B · 3B active | Q3_K_M | 17.2 GB |
| Ornith1.5 35B · 3B active | Q3_K_M | 17.2 GB |
| Qwen3.5 35B · 3B active | Q3_K_M | 17.2 GB |
| Agents-A1 35B · 3B active | Q3_K_M | 17.2 GB |
| Nex-N2.5-mini 35B · 3B active | Q3_K_M | 17.2 GB |
| KAT-Coder V2.5-Dev 35B · 3B active | Q3_K_M | 17.2 GB |
| Qwen2.5 32B | Q3_K_M | 17.9 GB |
| Qwen3 32B | Q3_K_M | 17.9 GB |
| Gemma 4 31B | Q4_K_M | 19.8 GB |
| Muse 30B | Q4_K_M | 18.8 GB |
| Qwen3 30B · 3B active | Q3_K_M | 17 GB |
| Nemotron 3.5 Lightning 30B · 3B active | Q4_K_M | 18.7 GB |
And 44 smaller models. The calculator lists them all for your card.
What people run on 24GB
Model sizes reported on setups with more than 16 GB and up to 24 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.827B Alibaba · mostly Q4_K_M Typical 35 t/s (16 runs) · with speculation 98 t/s (31) | 101 | 35 16 runs, 7 devices | 98 31 runs |
| 2 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly UD-Q4_K_XL No plain runs | 27 | no plain runs | none |
| 3 | Qwen3.627B Alibaba · mostly Q4_K_M Typical 42 t/s (3 runs) · with speculation 50 t/s (6) | 19 | 42 3 runs, 2 devices | 50 6 runs |
| 4 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly 2-bit dynamic No plain runs | 15 | no plain runs | none |
| 5 | Qwen3.635B · 3B active Alibaba · mostly IQ4 Typical 85 t/s (6 runs) · with speculation 52 t/s (1) | 14 | 85 6 runs, 3 devices | 52 1 run |
| 6 | Gemma 426B · 4B active Google DeepMind · mostly Q4_0 Typical 140 t/s (4 runs) | 6 | 140 4 runs, 2 devices | none |
| 7 | Muse Glimmer30B Meta · mostly IQ4_XS Typical 30 t/s (1 run) · with speculation 89 t/s (2) | 4 | 30 1 run | 89 2 runs |
| 8 | Gemma 431B Google DeepMind · mostly Q4_K_S Typical 33 t/s (2 runs) | 3 | 33 2 runs, 2 devices | none |
| 9 | Qwen3.8 Swift-1.527B Alibaba · mostly IQ4_XS Typical 67 t/s (1 run) · with speculation 104 t/s (1) | 2 | 67 1 run | 104 1 run |
| 10 | Qwen2.57B Alibaba · mostly GPTQ-Int4 Typical 66 t/s (2 runs) | 2 | 66 2 runs, 2 devices | none |
| 11 | Gemma 4 E4B8B Google DeepMind · mostly FP16 Typical 58 t/s (1 run) | 2 | 58 1 run | none |
| 12 | Nemotron 3.5 Lightning30B · 3B active NVIDIA · mostly Q4_0 No plain runs | 2 | no plain runs | none |
| 13 | Qwen3.8 Swift-1.5-Uncensored27B Alibaba · mostly IQ4_XS No plain runs · with speculation 68 t/s (1) | 2 | no plain runs | 68 1 run |
| 14 | Qwen2.5 Coder-7B7B Alibaba · mostly Q4_K_M Typical 159 t/s (1 run) | 1 | 159 1 run | none |
| 15 | DeepSeek-Coder-V2-Lite16B · 2.4B active mostly 4bit Typical 126 t/s (1 run) | 1 | 126 1 run | none |
| 16 | Qwen330B · 3B active Alibaba · mostly IQ3_XXS Typical 74 t/s (1 run) | 1 | 74 1 run | none |
| 17 | Qwen3.8 Escha-W227B Alibaba · mostly escha 2-bit Typical 67 t/s (1 run) | 1 | 67 1 run | none |
| 18 | Qwen3.59B Alibaba · mostly BF16 Typical 41 t/s (1 run) | 1 | 41 1 run | none |
| 19 | Qwen2.5 Coder-32B32B Alibaba · mostly Q4_K_M Typical 28 t/s (1 run) | 1 | 28 1 run | none |
| 20 | Qwen38B Alibaba · mostly Q4_K_M Typical 25 t/s (1 run) | 1 | 25 1 run | none |
| 21 | Qwen314B Alibaba · mostly 4-bit Typical 12 t/s (1 run) | 1 | 12 1 run | none |
| 22 | Qwen3.535B · 3B active Alibaba · mostly Q4_K_XL Typical 8.3 t/s (1 run) | 1 | 8.3 1 run | none |
| 23 | Qwen3 Heretic14B Alibaba · mostly INT4_SYM Typical 3.6 t/s (1 run) | 1 | 3.6 1 run | none |
| 24 | Bonsai 2 Ternary27B PrismML · mostly Q4/Q5 No plain runs · with speculation 188 t/s (1) | 1 | no plain runs | 188 1 run |
| 25 | DeepSeek V4 Flash Vision-Exp285B · 13B active DeepSeek · mostly FP4 No plain runs | 1 | no plain runs | none |
| 26 | DeepSeek V4.1 Flash552B · 16B active mostly 4-bit No plain runs | 1 | no plain runs | none |
| 27 | GLM-4.5-Air106B · 12B active No plain runs | 1 | no plain runs | none |
| 28 | GLM-5.2744B · 40B active Zhipu AI · mostly UD-IQ2_M No plain runs | 1 | no plain runs | none |
| 29 | GLM-5.3 Flash320B · 18B active Zhipu AI · mostly UD-Q4_K_XL No plain runs | 1 | no plain runs | none |
| 30 | Gemma 4 Uncensored-HauhauCS-Balanced26B · 4B active Google DeepMind · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 31 | Ornith1.5 BigBang35B · 3B active mostly Q4_K_M No plain runs · with speculation 169 t/s (1) | 1 | no plain runs | 169 1 run |
| 32 | Qwen2.532B Alibaba · mostly AWQ No plain runs | 1 | no plain runs | none |
| 33 | Qwen3235B · 22B active Alibaba · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 34 | Qwen3-Next80B · 3B active Alibaba · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 35 | Qwen3.5 Mica v0.14B Alibaba · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 36 | Qwen3.54B Alibaba No plain runs | 1 | no plain runs | none |
| 37 | Qwen3.6 Heretic-v227B Alibaba · mostly Q4_K_M No plain runs · with speculation 80 t/s (1) | 1 | no plain runs | 80 1 run |
| 38 | Qwen3.8 Swift-1.5-HyperQwen27B Alibaba · mostly W4A16 No plain runs | 1 | no plain runs | none |
| 39 | Qwen3.8 Swift27B Alibaba · mostly IQ4_XS No plain runs | 1 | no plain runs | none |
| 40 | Qwen3.8 CODER27B Alibaba · mostly IQ4_XS No plain runs | 1 | no plain runs | none |
| 41 | Qwen3.8 Huihui-Abliterated27B Alibaba · mostly Q4_K_S No plain runs | 1 | no plain runs | none |
| 42 | Qwen3.8 Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO27B Alibaba · mostly Q8_0 No plain runs | 1 | no plain runs | none |
| 43 | Qwen3.8 heretic-ara27B Alibaba · mostly Q5_K_M No plain runs | 1 | no plain runs | none |
Frequently asked
What do people run on 24GB of VRAM?
On setups with more than 16 GB and up to 24 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and Qwen3.6 27B. The table lists each with the quant most people used and the typical speed.
Does a Mac with 24GB count as 24GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 24GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.