llamaperf

Best local LLMs for 96GB VRAM

How far 96GB of graphics memory goes for local models on cards like NVIDIA RTX Pro 6000 Blackwell or NVIDIA RTX PRO 6000 Max-Q: which models fit, and what people actually run on it.

Is 96GB of VRAM enough?

96GB holds a dense model of up to about 137B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.8 125B · 6B active, which needs about 72.3 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.

The calculator checks your exact card, including longer contexts and offloading.

Models that fit in 96GB

Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.

And 53 smaller models. The calculator lists them all for your card.

What people run on 96GB

Model sizes reported on setups with more than 64 GB and up to 96 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.

#ModelReports
1Qwen3.8 Flash-Next125B · 6B active
Alibaba · mostly NVFP4
Typical 102 t/s (5 runs) · with speculation 43 t/s (4)
30
2Qwen3.827B
Alibaba · mostly FP8
Typical 18 t/s (3 runs) · with speculation 98 t/s (8)
23
3DeepSeek V4 Flash284B · 13B active
DeepSeek · mostly Q4_K_XL
Typical 44 t/s (3 runs)
14
4Qwen3.627B
Alibaba · mostly 8bit
Typical 7.4 t/s (1 run) · with speculation 28 t/s (3)
7
5GLM-5.3 Flash320B · 18B active
Zhipu AI · mostly 3
No plain runs
3
6Qwen3.635B · 3B active
Alibaba · mostly Q4_K_M
No plain runs
3
7Gemma 431B
Google DeepMind · mostly NVFP4
Typical 51 t/s (1 run) · with speculation 125 t/s (1)
2
8Muse Glimmer30B
Meta · mostly BF16
Typical 35 t/s (1 run) · with speculation 57 t/s (1)
2
9DeepSeek V4 Pro1600B · 49B active
DeepSeek
No plain runs
2
10DeepSeek V4.1 Flash552B · 16B active
mostly FP4
No plain runs
2
11Qwen330B · 3B active
Alibaba · mostly Q4_K_M
Typical 59 t/s (1 run)
1
12Qwen38B
Alibaba · mostly Q4
Typical 43 t/s (1 run)
1
13GLM-4.7 Flash30B
Zhipu AI
No plain runs · with speculation 168 t/s (1)
1
14GLM-5.1744B · 40B active
Zhipu AI · mostly NVFP4
No plain runs
1
15Llama 3.21B
Meta
No plain runs
1
16Mimo 2.5310B · 15B active
Xiaomi
No plain runs
1
17Mimo 2.6 Flash-RL309B · 15B active
Xiaomi · mostly 2.20 bpw
No plain runs
1
18Qwen3.5 NuExtract34B
Alibaba
No plain runs
1
19Qwen3.6 abliterated27B
Alibaba · mostly BF16
No plain runs
1
20Qwen3.6 Ornith-Agents-A1-3.635B · 3B active
Alibaba · mostly Q4_K_M
No plain runs
1
21Qwen3.8 Swift-1.527B
Alibaba · mostly 4.7bpw
No plain runs
1
22Tencent-HY3295B · 21B active
Tencent · mostly Q6_K
No plain runs
1

Frequently asked

What do people run on 96GB of VRAM?

On setups with more than 64 GB and up to 96 GB of memory, the model people report most is Qwen3.8 125B · 6B active, followed by Qwen3.8 27B and DeepSeek V4 Flash 284B · 13B active. The table lists each with the quant most people used and the typical speed.

Does a Mac with 96GB count as 96GB of VRAM?

No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 96GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.