llamaperf

Mimo 2.6 VRAM requirements

Memory needed to run Mimo 2.6 at every published size and quant, with an 8k context. Each figure is weights plus KV cache plus runtime buffers, from the same memory model as the VRAM calculator.

Total memory by size and quant

GB at an 8k context. Bold is the Q4_K_M column, the usual starting point.

SizeQ8_0Q6_KQ5_K_MQ4_K_MQ3_K_M
9B119.38.27.05.9
309B · 15B active311253214176137
1020B · 42B active1022831703576448

KV cache by context length

GB the context alone adds on top of the weights. Add this to a weights figure to size a longer session.

SizeWeights at Q4_K_M4k ctx16k ctx32k ctx128k ctx
9B5.10.10.40.83.0
309B · 15B active1740.10.40.83.0
1020B · 42B active5740.20.91.76.8

What each size fits on at Q4_K_M

Hardware whose memory holds the model with headroom, smallest pool first. Fits means under 85% of the pool at an 8k context. A Mac's pool is what macOS lets the GPU use - two thirds of its memory up to 32 GB, three quarters from 36 GB - not the whole of it.

Frequently asked

How much VRAM does Mimo 2.6 need?

From about 7.0 GB for the 9B model to about 576 GB for the 1020B · 42B active model, at Q4_K_M with an 8k context and including the KV cache and runtime buffers. Q8 needs more and Q3 less; the table above lists every size and rung.

Can I run Mimo 2.6 on a 16GB GPU?

Yes. The 9B model needs about 7.0 GB at Q4_K_M with an 8k context, which fits a 16 GB card with headroom. Larger sizes need a smaller quant or a bigger card.

Can I run Mimo 2.6 on a 24GB GPU such as an RTX 3090 or 4090?

Yes. The 9B model needs about 7.0 GB at Q4_K_M with an 8k context, which fits a 24 GB card with headroom. Larger sizes need a smaller quant or more memory.

How are these numbers calculated?

Weights at the quant's bits per weight, plus a KV cache sized from the model's attention design for the chosen context, plus compute buffers and a fixed runtime allowance. It is the same memory model the VRAM calculator uses. A family whose attention profile is not on file is sized as plain grouped-query attention, which errs towards needing more memory.