Best local LLMs for 128GB VRAM
How far 128GB of graphics memory goes for local models on cards like AMD Instinct MI250 or AMD Instinct MI250X 128GB: which models fit, and what people actually run on it.
Is 128GB of VRAM enough?
128GB holds a dense model of up to about 186B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.8 177B, which needs about 101.9 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 128GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| Qwen3 235B · 22B active | Q3_K_M | 106.7 GB |
| Minimax M2.1 230B · 10B active | Q3_K_M | 104.5 GB |
| StepFun 3.7 198B · 11B active | Q3_K_M | 88.9 GB |
| Qwen3.8 177B | Q4_K_M | 101.9 GB |
| DeepSeek V4 Flash 150B | Q5_K_M | 105.3 GB |
| Qwen3.8 125B · 6B active | Q6_K | 103.5 GB |
| Ling-3.0 124B · 5.1B active | Q6_K | 102.6 GB |
| Qwen3.5 122B · 10B active | Q6_K | 101.7 GB |
| GLM-4.5-Air 106B · 12B active | Q6_K | 89.4 GB |
| Qwen3-Coder-Next 80B · 3B active | Q8_0 | 81.9 GB |
| Qwen3-Next 80B · 3B active | Q8_0 | 82.5 GB |
| Qwen2.5 72B | Q8_0 | 76.5 GB |
And 56 smaller models. The calculator lists them all for your card.
What people run on 128GB
Model sizes reported on setups with more than 96 GB and up to 128 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly NVFP4 Typical 38 t/s (12 runs) · with speculation 42 t/s (18) | 58 | 38 12 runs, 5 devices | 42 18 runs |
| 2 | Qwen3.827B Alibaba · mostly NVFP4 Typical 14 t/s (5 runs) · with speculation 40 t/s (16) | 30 | 14 5 runs, 2 devices | 40 16 runs |
| 3 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly UD-IQ3_XXS Typical 35 t/s (1 run) · with speculation 29 t/s (3) | 16 | 35 1 run | 29 3 runs |
| 4 | Qwen3.635B · 3B active Alibaba · mostly NVFP4 Typical 62 t/s (6 runs) · with speculation 102 t/s (4) | 12 | 62 6 runs, 4 devices | 102 4 runs |
| 5 | Qwen3.627B Alibaba · mostly Q8 Typical 15 t/s (2 runs) · with speculation 21 t/s (3) | 11 | 15 2 runs, 2 devices | 21 3 runs |
| 6 | DeepSeek V4.1 Flash552B · 16B active mostly Q2 No plain runs | 4 | no plain runs | none |
| 7 | Ornith1.535B · 3B active mostly NVFP4 Typical 71 t/s (2 runs) · with speculation 79 t/s (1) | 3 | 71 2 runs, 2 devices | 79 1 run |
| 8 | Ling-3.0 Flash124B · 5.1B active Ant Group · mostly Q5_K_M Typical 39 t/s (3 runs) | 3 | 39 3 runs | none |
| 9 | Nex-N2.5-mini35B · 3B active mostly MLX-4bit Typical 105 t/s (2 runs) | 2 | 105 2 runs, 2 devices | none |
| 10 | Qwen330B · 3B active Alibaba · mostly 4-bit Typical 128 t/s (1 run) | 1 | 128 1 run | none |
| 11 | LFM2.52.6B Liquid AI · mostly Q4_K_M Typical 113 t/s (1 run) | 1 | 113 1 run | none |
| 12 | Qwen3.535B · 3B active Alibaba · mostly 4-bit Typical 91 t/s (1 run) | 1 | 91 1 run | none |
| 13 | Gemma 426B · 4B active Google DeepMind · mostly Q4_K_M Typical 70 t/s (1 run) | 1 | 70 1 run | none |
| 14 | Qwen3-Coder-Next80B · 3B active mostly UD-Q6_K_XL Typical 37 t/s (1 run) | 1 | 37 1 run | none |
| 15 | StepFun 3.7 Flash198B · 11B active StepFun · mostly Q4_K_S Typical 34 t/s (1 run) | 1 | 34 1 run | none |
| 16 | Tencent-HY3295B · 21B active Tencent · mostly UD128 Typical 32 t/s (1 run) | 1 | 32 1 run | none |
| 17 | Mimo 2.6 Flash-RL309B · 15B active Xiaomi · mostly IQ2_M Typical 26 t/s (1 run) | 1 | 26 1 run | none |
| 18 | Muse Glimmer30B Meta · mostly UD-Q2_K_XL Typical 25 t/s (1 run) | 1 | 25 1 run | none |
| 19 | GLM-5.3 Flash320B · 18B active Zhipu AI · mostly EXL3 No plain runs | 1 | no plain runs | none |
| 20 | Gemma 431B Google DeepMind No plain runs · with speculation 7.5 t/s (1) | 1 | no plain runs | 7.5 1 run |
| 21 | Kimi K32800B · 104B active Moonshot AI · mostly 4-bit No plain runs | 1 | no plain runs | none |
| 22 | Nemotron 3.5 Lightning30B · 3B active NVIDIA · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 23 | Qwen3.5122B · 10B active Alibaba · mostly WinterMix58 No plain runs | 1 | no plain runs | none |
| 24 | Qwen3.8 Swift27B Alibaba · mostly 4bit No plain runs · with speculation 88 t/s (1) | 1 | no plain runs | 88 1 run |
Frequently asked
What do people run on 128GB of VRAM?
On setups with more than 96 GB and up to 128 GB of memory, the model people report most is Qwen3.8 125B · 6B active, followed by Qwen3.8 27B and DeepSeek V4 Flash 284B · 13B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 128GB count as 128GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 128GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.