llamaperf

Best local LLMs for 128GB VRAM

How far 128GB of graphics memory goes for local models on cards like AMD Instinct MI250 or AMD Instinct MI250X 128GB: which models fit, and what people actually run on it.

Is 128GB of VRAM enough?

128GB holds a dense model of up to about 186B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.8 177B, which needs about 101.9 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.

The calculator checks your exact card, including longer contexts and offloading.

Models that fit in 128GB

Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.

And 56 smaller models. The calculator lists them all for your card.

What people run on 128GB

Model sizes reported on setups with more than 96 GB and up to 128 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.

#ModelReports
1Qwen3.8 Flash-Next125B · 6B active
Alibaba · mostly NVFP4
Typical 38 t/s (12 runs) · with speculation 42 t/s (18)
58
2Qwen3.827B
Alibaba · mostly NVFP4
Typical 14 t/s (5 runs) · with speculation 40 t/s (16)
30
3DeepSeek V4 Flash284B · 13B active
DeepSeek · mostly UD-IQ3_XXS
Typical 35 t/s (1 run) · with speculation 29 t/s (3)
16
4Qwen3.635B · 3B active
Alibaba · mostly NVFP4
Typical 62 t/s (6 runs) · with speculation 102 t/s (4)
12
5Qwen3.627B
Alibaba · mostly Q8
Typical 15 t/s (2 runs) · with speculation 21 t/s (3)
11
6DeepSeek V4.1 Flash552B · 16B active
mostly Q2
No plain runs
4
7Ornith1.535B · 3B active
mostly NVFP4
Typical 71 t/s (2 runs) · with speculation 79 t/s (1)
3
8Ling-3.0 Flash124B · 5.1B active
Ant Group · mostly Q5_K_M
Typical 39 t/s (3 runs)
3
9Nex-N2.5-mini35B · 3B active
mostly MLX-4bit
Typical 105 t/s (2 runs)
2
10Qwen330B · 3B active
Alibaba · mostly 4-bit
Typical 128 t/s (1 run)
1
11LFM2.52.6B
Liquid AI · mostly Q4_K_M
Typical 113 t/s (1 run)
1
12Qwen3.535B · 3B active
Alibaba · mostly 4-bit
Typical 91 t/s (1 run)
1
13Gemma 426B · 4B active
Google DeepMind · mostly Q4_K_M
Typical 70 t/s (1 run)
1
14Qwen3-Coder-Next80B · 3B active
mostly UD-Q6_K_XL
Typical 37 t/s (1 run)
1
15StepFun 3.7 Flash198B · 11B active
StepFun · mostly Q4_K_S
Typical 34 t/s (1 run)
1
16Tencent-HY3295B · 21B active
Tencent · mostly UD128
Typical 32 t/s (1 run)
1
17Mimo 2.6 Flash-RL309B · 15B active
Xiaomi · mostly IQ2_M
Typical 26 t/s (1 run)
1
18Muse Glimmer30B
Meta · mostly UD-Q2_K_XL
Typical 25 t/s (1 run)
1
19GLM-5.3 Flash320B · 18B active
Zhipu AI · mostly EXL3
No plain runs
1
20Gemma 431B
Google DeepMind
No plain runs · with speculation 7.5 t/s (1)
1
21Kimi K32800B · 104B active
Moonshot AI · mostly 4-bit
No plain runs
1
22Nemotron 3.5 Lightning30B · 3B active
NVIDIA · mostly Q4_K_M
No plain runs
1
23Qwen3.5122B · 10B active
Alibaba · mostly WinterMix58
No plain runs
1
24Qwen3.8 Swift27B
Alibaba · mostly 4bit
No plain runs · with speculation 88 t/s (1)
1

Frequently asked

What do people run on 128GB of VRAM?

On setups with more than 96 GB and up to 128 GB of memory, the model people report most is Qwen3.8 125B · 6B active, followed by Qwen3.8 27B and DeepSeek V4 Flash 284B · 13B active. The table lists each with the quant most people used and the typical speed.

Does a Mac with 128GB count as 128GB of VRAM?

No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 128GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.