Best local LLMs for 48GB VRAM
How far 48GB of graphics memory goes for local models on cards like NVIDIA A40 48GB, NVIDIA RTX 6000 Ada, NVIDIA RTX A6000 48GB or AMD Radeon Pro W7800 48GB: which models fit, and what people actually run on it.
Is 48GB of VRAM enough?
48GB holds a dense model of up to about 64B parameters at Q4_K_M with an 8K context, by the calculator's memory estimate. The largest model people report here that fits comfortably is Qwen3.6 35B · 3B active, which needs about 36.9 GB at Q8_0 with an 8K context. Anything bigger needs a smaller quant, part of the model in system RAM, or more memory, and a longer context needs more room on top.
The calculator checks your exact card, including longer contexts and offloading.
Models that fit in 48GB
Models people report on llamaperf, largest first, at the best quant whose memory with an 8K context stays within the calculator's comfortable band (85% of the card). Estimated from each model's shape.
| Model | Best quant | Needs |
|---|---|---|
| Qwen3-Coder-Next 80B · 3B active | Q3_K_M | 36.9 GB |
| Qwen3-Next 80B · 3B active | Q3_K_M | 37.5 GB |
| Qwen2.5 72B | Q3_K_M | 36 GB |
| Qwen3.6 35B · 3B active | Q8_0 | 36.9 GB |
| Ornith1.5 35B · 3B active | Q8_0 | 36.9 GB |
| Qwen3.5 35B · 3B active | Q8_0 | 36.9 GB |
| Agents-A1 35B · 3B active | Q8_0 | 36.9 GB |
| Nex-N2.5-mini 35B · 3B active | Q8_0 | 36.9 GB |
| KAT-Coder V2.5-Dev 35B · 3B active | Q8_0 | 36.9 GB |
| Qwen2.5 32B | Q8_0 | 35.9 GB |
| Qwen3 32B | Q8_0 | 35.9 GB |
| Gemma 4 31B | Q8_0 | 33.4 GB |
And 47 smaller models. The calculator lists them all for your card.
What people run on 48GB
Model sizes reported on setups with more than 32 GB and up to 48 GB of memory, most reported first. Typical speed is the median of plain runs, and speculative runs are counted apart.
| # | Model | Reports | Typical t/s | With speculation |
|---|---|---|---|---|
| 1 | Qwen3.827B Alibaba · mostly Q4 No plain runs · with speculation 20 t/s (5) | 24 | no plain runs | 20 5 runs |
| 2 | Qwen3.8 Flash-Next125B · 6B active Alibaba · mostly UD-Q4_K_XL Typical 90 t/s (1 run) | 21 | 90 1 run | none |
| 3 | DeepSeek V4 Flash284B · 13B active DeepSeek · mostly Q8_K_XL No plain runs | 7 | no plain runs | none |
| 4 | Qwen3.627B Alibaba · mostly 4-bit No plain runs · with speculation 49 t/s (2) | 4 | no plain runs | 49 2 runs |
| 5 | Qwen3.54B Alibaba · mostly Q4_K_M Typical 85 t/s (1 run) | 2 | 85 1 run | none |
| 6 | Qwen3.635B · 3B active Alibaba · mostly 4bit Typical 21 t/s (1 run) · with speculation 110 t/s (1) | 2 | 21 1 run | 110 1 run |
| 7 | Qwen3.8 Uncensored27B Alibaba · mostly 4-bit No plain runs · with speculation 92 t/s (1) | 2 | no plain runs | 92 1 run |
| 8 | DeepSeek V4 Flash 9B distill9B DeepSeek · mostly Q4_K_M Typical 44 t/s (1 run) | 1 | 44 1 run | none |
| 9 | Gemma 4 E2B5.1B Google DeepMind · mostly bf16 Typical 17 t/s (1 run) | 1 | 17 1 run | none |
| 10 | Muse Glimmer30B Meta · mostly 8-bit Typical 8.2 t/s (1 run) | 1 | 8.2 1 run | none |
| 11 | DeepSeek R1671B · 37B active DeepSeek · mostly Q3 No plain runs | 1 | no plain runs | none |
| 12 | Qwen2.57B Alibaba No plain runs | 1 | no plain runs | none |
| 13 | Qwen3.8 Swift27B Alibaba · mostly Q4_K_M No plain runs | 1 | no plain runs | none |
| 14 | Qwen3.8 Swift-1.5-125B-A6B125B · 6B active Alibaba · mostly IQ3-XXS No plain runs | 1 | no plain runs | none |
| 15 | Qwen3.8 Flash-Next Uncensored-177B177B Alibaba · mostly IQ3_XXS No plain runs | 1 | no plain runs | none |
Frequently asked
What do people run on 48GB of VRAM?
On setups with more than 32 GB and up to 48 GB of memory, the model people report most is Qwen3.8 27B, followed by Qwen3.8 125B · 6B active and DeepSeek V4 Flash 284B · 13B active. The table lists each with the quant most people used and the typical speed.
Does a Mac with 48GB count as 48GB of VRAM?
No. macOS keeps part of a Mac's unified memory back from the GPU, so a Mac with 48GB can use less than that for a model. The calculator's Mac tab grades every Mac against its real limit.