gpt-oss VRAM requirements
Memory needed to run gpt-oss at every published size and quant, with an 8k context. Each figure is weights plus KV cache plus runtime buffers, from the same memory model as the VRAM calculator. This family's attention design is not on file, so it is sized as plain grouped-query attention, which errs towards needing more.
Total memory by size and quant
GB at an 8k context. Bold is the Q4_K_M column, the usual starting point.
| Size | Q8_0 | Q6_K | Q5_K_M | Q4_K_M | Q3_K_M |
|---|---|---|---|---|---|
| 21B · 3.6B active | 25 | 21 | 18 | 16 | 13 |
| 117B · 5.1B active | 122 | 100 | 85 | 71 | 56 |
KV cache by context length
GB the context alone adds on top of the weights. Add this to a weights figure to size a longer session.
| Size | Weights at Q4_K_M | 4k ctx | 16k ctx | 32k ctx | 128k ctx |
|---|---|---|---|---|---|
| 21B · 3.6B active | 12 | 1.1 | 4.3 | 8.6 | 34 |
| 117B · 5.1B active | 66 | 1.5 | 5.9 | 12 | 47 |
What each size fits on at Q4_K_M
Hardware whose memory holds the model with headroom, smallest pool first. Fits means under 85% of the pool at an 8k context. A Mac's pool is what macOS lets the GPU use - two thirds of its memory up to 32 GB, three quarters from 36 GB - not the whole of it.
- 21B · 3.6B activeneeds about 16 GB
Smallest: AMD RX 7900 XT (20 GB), NVIDIA RTX 3080 20GB (20 GB), NVIDIA RTX 4000 Ada (20 GB), NVIDIA RTX 2080 Ti 22GB (modded) (22 GB), and 60 more.
On a Mac: 32 GB or more, with macOS's default GPU limit.
- 117B · 5.1B activeneeds about 71 GB
Smallest: NVIDIA H100 NVL (94 GB), NVIDIA RTX Pro 6000 Blackwell (96 GB), NVIDIA RTX PRO 6000 Max-Q (96 GB), AMD Instinct MI250 (128 GB), and 13 more.
On a Mac: 128 GB or more, with macOS's default GPU limit.
Frequently asked
How much VRAM does gpt-oss need?
From about 16 GB for the 21B · 3.6B active model to about 71 GB for the 117B · 5.1B active model, at Q4_K_M with an 8k context and including the KV cache and runtime buffers. Q8 needs more and Q3 less; the table above lists every size and rung.
Can I run gpt-oss on a 16GB GPU?
Not comfortably. Even the 21B · 3.6B active model needs about 16 GB at Q4_K_M with an 8k context, which is more than a 16 GB card holds with headroom. A lower quant or CPU offloading can still get it running, at a cost in quality or speed.
Can I run gpt-oss on a 24GB GPU such as an RTX 3090 or 4090?
Yes. The 21B · 3.6B active model needs about 16 GB at Q4_K_M with an 8k context, which fits a 24 GB card with headroom. Larger sizes need a smaller quant or more memory.
How are these numbers calculated?
Weights at the quant's bits per weight, plus a KV cache sized from the model's attention design for the chosen context, plus compute buffers and a fixed runtime allowance. It is the same memory model the VRAM calculator uses. A family whose attention profile is not on file is sized as plain grouped-query attention, which errs towards needing more memory.