Best local LLMs for 64GB VRAM
How far 64GB of graphics memory goes for local models on cards like NVIDIA CMP 170HX 64GB (unlocked) or AMD Instinct MI210: which models fit, and what people actually run on it.
Is 64GB of VRAM enough?
64GB holds a dense model of up to about 87B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3-Coder-Next 80B · 3B active, which needs about 46.9 GB at Q4_K_M with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 64GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| GLM-4.5-Air 106B · 12B active | Q3_K_M | 49.7 GB |
| Qwen3-Coder-Next 80B · 3B active | Q4_K_M | 46.9 GB |
| Qwen3-Next 80B · 3B active | Q4_K_M | 47.5 GB |
| Qwen2.5 72B | Q5_K_M | 54 GB |
| Qwen3.6 35B · 3B active | Q8_0 | 36.9 GB |
| Ornith1.5 35B · 3B active | Q8_0 | 36.9 GB |
| Qwen3.5 35B · 3B active | Q8_0 | 36.9 GB |
| Agents-A1 35B · 3B active | Q8_0 | 36.9 GB |
| Nex-N2.5-mini 35B · 3B active | Q8_0 | 36.9 GB |
| KAT-Coder V2.5-Dev 35B · 3B active | Q8_0 | 36.9 GB |
| Qwen2.5 32B | Q8_0 | 35.9 GB |
| Qwen3 32B | Q8_0 | 35.9 GB |
And 48 smaller models. The calculator lists them all for your card.
What people run on 64GB
Model sizes reported on setups with more than 48 GB and up to 64 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.827B Alibaba · mostly Q8_0 Typical 35 t/s (2 runs) · with speculation 34 t/s (8) | 20 | 35 2 runs, 2 devices | 34 8 runs |
| 2 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly IQ4_XS Typical 31 t/s (2 runs) · with speculation 44 t/s (1) | 20 | 31 2 runs, 2 devices | 44 1 run |
| 3 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly IQ3-XXS No plain runs | 9 | no plain runs | none |
| 4 | Qwen3.635B · 3B active Alibaba · mostly 4-bit Typical 71 t/s (1 run) · with speculation 121 t/s (3) | 4 | 71 1 run | 121 3 runs |
| 5 | Qwen3.627B Alibaba · mostly 4-bit Typical 32 t/s (1 run) · with speculation 63 t/s (1) | 3 | 32 1 run | 63 1 run |
| 6 | Qwen3.8 Swift-1.527B Alibaba · mostly FP8 No plain runs · with speculation 35 t/s (1) | 3 | no plain runs | 35 1 run |
| 7 | DeepSeek V4.1 Flash552B · 16B active No plain runs | 2 | no plain runs | none |
| 8 | Qwen3-Next80B · 3B active Alibaba · mostly W4A16 Typical 116 t/s (1 run) | 1 | 116 1 run | none |
| 9 | Gemma 426B · 4B active Google DeepMind Typical 97 t/s (1 run) | 1 | 97 1 run | none |
| 10 | Phi-3.5 mini3.8B Microsoft · mostly 4bit Typical 72 t/s (1 run) | 1 | 72 1 run | none |
| 11 | Qwen3.54B Alibaba · mostly Q4_K_M Typical 70 t/s (1 run) | 1 | 70 1 run | none |
| 12 | Qwen3.535B · 3B active Alibaba · mostly GPTQ-Int4 Typical 31 t/s (1 run) | 1 | 31 1 run | none |
| 13 | DeepSeek V4 Flash284B DeepSeek · mostly q2 No plain runs | 1 | no plain runs | none |
| 14 | GLM-5.3 Flash320B · 18B active Zhipu AI · mostly IQ4_XS No plain runs | 1 | no plain runs | none |
| 15 | Kimi K32800B · 104B active Moonshot AI No plain runs | 1 | no plain runs | none |
| 16 | Muse Glimmer30B Meta · mostly Q6_K_XL No plain runs | 1 | no plain runs | none |
| 17 | Qwen3.8 Uncensored27B Alibaba · mostly Q8_0 No plain runs | 1 | no plain runs | none |
| 18 | Qwen3.8 Splash-HQ27B Alibaba · mostly Q8 No plain runs · with speculation 37 t/s (1) | 1 | no plain runs | 37 1 run |
| 19 | Qwen3.82400B · 95B active Alibaba · mostly Q1_0 No plain runs | 1 | no plain runs | none |
Frequently asked
What do people run on 64GB of VRAM?
On setups with more than 48 GB and up to 64 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and DeepSeek V4 Flash 284B · 13B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 64GB count as 64GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 64GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.