Best local LLMs for 96GB VRAM
How far 96GB of graphics memory goes for local models on cards like NVIDIA RTX Pro 6000 Blackwell or NVIDIA RTX PRO 6000 Max-Q: which models fit, and what people actually run on it.
Is 96GB of VRAM enough?
96GB holds a dense model of up to about 137B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.8 125B · 6B active, which needs about 72.3 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 96GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| Qwen3.8 177B | Q3_K_M | 79.8 GB |
| DeepSeek V4 Flash 150B | Q3_K_M | 67.8 GB |
| Qwen3.8 125B · 6B active | Q4_K_M | 72.3 GB |
| Ling-3.0 124B · 5.1B active | Q4_K_M | 71.6 GB |
| Qwen3.5 122B · 10B active | Q4_K_M | 71.2 GB |
| GLM-4.5-Air 106B · 12B active | Q5_K_M | 76.2 GB |
| Qwen3-Coder-Next 80B · 3B active | Q6_K | 66.9 GB |
| Qwen3-Next 80B · 3B active | Q6_K | 67.5 GB |
| Qwen2.5 72B | Q8_0 | 76.5 GB |
| Qwen3.6 35B · 3B active | Q8_0 | 36.9 GB |
| Ornith1.5 35B · 3B active | Q8_0 | 36.9 GB |
| Qwen3.5 35B · 3B active | Q8_0 | 36.9 GB |
And 53 smaller models. The calculator lists them all for your card.
What people run on 96GB
Model sizes reported on setups with more than 64 GB and up to 96 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly NVFP4 Typical 102 t/s (5 runs) · with speculation 43 t/s (4) | 30 | 102 5 runs | 43 4 runs |
| 2 | Qwen3.827B Alibaba · mostly FP8 Typical 18 t/s (3 runs) · with speculation 98 t/s (8) | 23 | 18 3 runs, 3 devices | 98 8 runs |
| 3 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly Q4_K_XL Typical 44 t/s (3 runs) | 14 | 44 3 runs, 3 devices | none |
| 4 | Qwen3.627B Alibaba · mostly 8bit Typical 7.4 t/s (1 run) · with speculation 28 t/s (3) | 7 | 7.4 1 run | 28 3 runs |
| 5 | GLM-5.3 Flash320B · 18B active Zhipu AI · mostly 3 No plain runs | 3 | no plain runs | none |
| 6 | Qwen3.635B · 3B active Alibaba · mostly Q4_K_M No plain runs | 3 | no plain runs | none |
| 7 | Gemma 431B Google DeepMind · mostly NVFP4 Typical 51 t/s (1 run) · with speculation 125 t/s (1) | 2 | 51 1 run | 125 1 run |
| 8 | Muse Glimmer30B Meta · mostly BF16 Typical 35 t/s (1 run) · with speculation 57 t/s (1) | 2 | 35 1 run | 57 1 run |
| 9 | DeepSeek V4 Pro1600B · 49B active DeepSeek No plain runs | 2 | no plain runs | none |
| 10 | DeepSeek V4.1 Flash552B · 16B active mostly FP4 No plain runs | 2 | no plain runs | none |
| 11 | Qwen330B · 3B active Alibaba · mostly Q4_K_M Typical 59 t/s (1 run) | 1 | 59 1 run | none |
| 12 | Qwen38B Alibaba · mostly Q4 Typical 43 t/s (1 run) | 1 | 43 1 run | none |
| 13 | GLM-4.7 Flash30B Zhipu AI No plain runs · with speculation 168 t/s (1) | 1 | no plain runs | 168 1 run |
| 14 | GLM-5.1744B · 40B active Zhipu AI · mostly NVFP4 No plain runs | 1 | no plain runs | none |
| 15 | Llama 3.21B Meta No plain runs | 1 | no plain runs | none |
| 16 | Mimo 2.5310B · 15B active Xiaomi No plain runs | 1 | no plain runs | none |
| 17 | Mimo 2.6 Flash-RL309B · 15B active Xiaomi · mostly 2.20 bpw No plain runs | 1 | no plain runs | none |
| 18 | Qwen3.5 NuExtract34B Alibaba No plain runs | 1 | no plain runs | none |
| 19 | Qwen3.6 abliterated27B Alibaba · mostly BF16 No plain runs | 1 | no plain runs | none |
| 20 | Qwen3.6 Ornith-Agents-A1-3.635B · 3B active Alibaba · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 21 | Qwen3.8 Swift-1.527B Alibaba · mostly 4.7bpw No plain runs | 1 | no plain runs | none |
| 22 | Tencent-HY3295B · 21B active Tencent · mostly Q6_K No plain runs | 1 | no plain runs | none |
Frequently asked
What do people run on 96GB of VRAM?
On setups with more than 64 GB and up to 96 GB of memory, the model people report most is Qwen3.8 125B · 6B active, followed by Qwen3.8 27B and DeepSeek V4 Flash 284B · 13B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 96GB count as 96GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 96GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.