Best local LLM for coding with 24GB VRAM
The models people report running for coding on setups with more than 16 GB and up to 24 GB of memory, such as RTX 3090, RTX 4090, RX 7900 XTX or L4. Each row shows the fastest coding run reported for that model in this band.
Ranked from 33 coding reports on llamaperf.
Ranked by community reports
| # | Model | Fastest run | Median t/s | Fastest t/s | Reports |
|---|---|---|---|---|---|
| 1 | Qwen3.8Alibaba | 27B · INT4 on RTX 3090 | 44 | 165 | 16 |
| 2 | Qwen3.6Alibaba | 35B · 3B active · ninfer quant on RTX 3090 | 50 | 171 | 11 |
| 3 | DeepSeek V4 FlashDeepSeek | 284B · 13B active · UD-Q3_K_M on RTX 4090 | 12 | 13 | 2 |
| 4 | MuseMeta | 30B · Q5_K_M | 85 | 85 | 1 |
| 5 | Gemma 4Google DeepMind | 26B · 4B active on RTX 4090 | 138 | 138 | 1 |
| 6 | Qwen2.5Alibaba | 32B · Q4_K_M on RTX 3090 | 28 | 28 | 1 |
| 7 | GLM-5.2Zhipu AI | Q1_S on RTX 3090 | 6.0 | 6.0 | 1 |
Frequently asked
What is the best local LLM for coding with 24GB of VRAM?
Ranked from community reports on setups with more than 16 GB and up to 24 GB of memory, Qwen3.8 has the strongest record, followed by Qwen3.6 and DeepSeek V4 Flash. The table shows the quant and GPU of each model's fastest coding run so you can copy a setup that is known to work.
What counts as a coding report?
A community performance report whose poster described using the model for coding: an editor assistant, an agent, or code generation. The memory band is the poster's reported VRAM, or the card's VRAM times the number of cards.
Which quant should I use for coding on 24GB?
Start from the quant in the fastest run column, which is one that is known to fit with room for a coding context. If you want a bigger model in the same memory, step down one quant rung; the VRAM calculator shows the exact memory at each rung for your card.
How we rank
Families are scored on report count, the fastest generation speed within this memory band, recency, and how complete the best report is. Reports come from r/LocalLLaMA and direct submissions. Nothing is sponsored.