llamaperf

Best local LLM for coding with 128GB VRAM

The models people report running for coding on setups with more than 96 GB and up to 128 GB of memory, such as AMD Strix Halo 128GB, DGX Spark, M5 Max 128GB or M4 Max 128GB. Each row shows the fastest coding run reported for that model in this band.

Ranked from 19 coding reports on llamaperf.

Ranked by community reports

#ModelFastest runMedian t/sFastest t/sReports
1Qwen3.8Alibaba125B · 6B active · oQ4e on M5 Max 128GB26558
2Qwen3.6Alibaba27B · Q4_K_M on AMD Strix Halo 128GB17216
3Nex-N2.5-miniMLX-4bit on M5 Max 128GB1341341
4DeepSeek V4 FlashDeepSeek284B · 13B active · UD-IQ3_XXS on DGX Spark24241
5Ling-3.0Ant GroupINT4 on DGX Spark41411
6Qwen3-Coder-NextUD-Q6_K_XL on AMD Strix Halo 128GB37371
7GLM-5.2Zhipu AINVFP4 on DGX Spark15151

Frequently asked

What is the best local LLM for coding with 128GB of VRAM?

Ranked from community reports on setups with more than 96 GB and up to 128 GB of memory, Qwen3.8 has the strongest record, followed by Qwen3.6 and Nex-N2.5-mini. The table shows the quant and GPU of each model's fastest coding run so you can copy a setup that is known to work.

What counts as a coding report?

A community performance report whose poster described using the model for coding: an editor assistant, an agent, or code generation. The memory band is the poster's reported VRAM, or the card's VRAM times the number of cards.

Which quant should I use for coding on 128GB?

Start from the quant in the fastest run column, which is one that is known to fit with room for a coding context. If you want a bigger model in the same memory, step down one quant rung; the VRAM calculator shows the exact memory at each rung for your card.

How we rank

Families are scored on report count, the fastest generation speed within this memory band, recency, and how complete the best report is. Reports come from r/LocalLLaMA and direct submissions. Nothing is sponsored.