Best local LLM for coding
The models people report running for coding on their own hardware, ranked from community reports. Pick your memory size first: a model that is the best choice on a 24 GB card is often not the best on 12 GB.
Ranked from 158 coding reports on llamaperf.
By VRAM
All memory sizes
| # | Model | Fastest run | Median t/s | Fastest t/s | Reports |
|---|---|---|---|---|---|
| 1 | Qwen3.8Alibaba | 27B · NVFP4 on RTX 5090 | 40 | 201 | 70 |
| 2 | Qwen3.6Alibaba | 35B · 3B active · ninfer quant on RTX 3090 | 48 | 171 | 45 |
| 3 | DeepSeek V4 FlashDeepSeek | 284B · 13B active | 16 | 180 | 16 |
| 4 | MuseMeta | 30B · UD-Q5_K_M on RTX 5090 | 85 | 253 | 7 |
| 5 | Gemma 4Google DeepMind | 26B · 4B active on RTX 4090 | 111 | 138 | 4 |
| 6 | Ornith1.5 | 9B · Q6_K | 30 | 36 | 2 |
| 7 | Qwen2.5Alibaba | 27B · Q6_K_XL on RTX 3090 | 49 | 70 | 2 |
| 8 | GLM-5.2Zhipu AI | NVFP4 on DGX Spark | 10 | 15 | 2 |
| 9 | Nex-N2.5-mini | MLX-4bit on M5 Max 128GB | 134 | 134 | 1 |
| 10 | Qwen3Alibaba | 8B · Q4 on M2 Max 96GB | 43 | 43 | 1 |
| 11 | Agents-A1InternScience | 35B · 3B active · ternary | 22 | 22 | 1 |
| 12 | GLM-5.3Zhipu AI | Q4 on M3 Ultra 256GB | 37 | 37 | 1 |
| 13 | Qwen3.5Alibaba | 9B · Q4 on D700 12GB | 11 | 11 | 1 |
| 14 | Ling-3.0Ant Group | INT4 on DGX Spark | 41 | 41 | 1 |
| 15 | Qwen3-Coder-Next | UD-Q6_K_XL on AMD Strix Halo 128GB | 37 | 37 | 1 |
| 16 | DeepSeek V4.1 Flash | oQ4e on M3 Ultra 512GB | 20 | 20 | 1 |
| 17 | KAT-Coder V2.5-DevKwaipilot | RTX 3060 12GB | 14 | 14 | 1 |
| 18 | Mimo 2.5Xiaomi | RTX Pro 6000 Blackwell | - | - | 1 |
Frequently asked
What is the best local LLM for coding?
It depends on how much memory you have, which is why this page is split by VRAM. Each tier ranks the models people report using for coding on that much memory, with the quant and GPU of the fastest run, so you can pick from setups like yours rather than from a single global list.
How are these rankings built?
From community performance reports tagged as coding use. Models are scored on how many reports they have, the fastest generation speed within the tier, how recent the reports are, and how complete the best report is. Nothing is sponsored.
Does a coding model need more VRAM than a chat model?
Not for the weights, but coding sessions carry long contexts: whole files, diffs and tool output. The KV cache for that context has to fit next to the weights, so leave more headroom than a chat use would need, or pick a quant one rung smaller.