llamaperf
← All articles

Quantization

What Q4_K_M means: reading GGUF quantization names

By llamaperf · · 6 min read

Quick answer

The number is roughly how many bits each weight gets. The rest names the recipe: K is llama.cpp's k-quant method, and the final S, M or L says how many of the most sensitive tensors get a bigger type. Q4_K_M lands near 4.9 bits per weight in practice, so an 8B model comes out around 4.6 GiB before you add any context.

How to read the name

Open any GGUF repository and you'll see a dozen files that differ only in a short code at the end: Q4_K_M, Q5_K_S, Q6_K, Q8_0, IQ4_XS. Each code is a recipe for how the model's weights were squeezed. Once you can read it, you can guess the file size and roughly how much quality you're trading before you download 20 GB.

The number and the letters

Q means quantized. The number after it is the base bit width for most of the weights, so Q4 stores them at about four bits and Q8 at about eight. The K means the file uses llama.cpp's k-quants, which group weights into blocks and super-blocks that each carry their own scale. The Hugging Face GGUF reference lists every type with its layout. Its Q4_K entry works out to 4.5 bits per weight and Q6_K to 6.5625, because the scales take up space too.

What S, M and L change

The trailing letter is small, medium or large. The original k-quants pull request spells out the recipes. Q4_K_S uses the 4-bit type for every tensor. Q4_K_M gives half of two sensitive tensor groups, the attention value weights and one feed-forward projection, the 6-bit type and keeps the rest at 4 bits. So M and L spend a little more space where errors hurt the most. llama.cpp has tuned these recipes since 2023, so treat the PR as the idea and your file's metadata as the final word.

From bits per weight to file size

File size is close to parameters times bits per weight, divided by eight. That's a handy check before you download anything.

Why Q4_K_M is bigger than 4 bits

The llama.cpp team publishes measured sizes for Llama 3.1 8B in the quantize tool's README. They show how far the real files sit from the name:

TypeBits per weightSize (GiB)
Q4_K_S4.66724.36
Q4_K_M4.89444.58
Q5_K_M5.70365.33
Q6_K6.56336.14
Q8_08.50087.95
F1616.000514.96

Q4_K_M comes out near 4.9 bits because of the block scales and the tensors it bumps up to 6 bits. The same README sums up why anyone bothers:

Quantization reduces the precision of model weights (e.g., from 32-bit floats to 4-bit integers), which shrinks the model's size and can speed up inference. From the llama.cpp quantize README

Estimate a download before you grab it

Say you're looking at a 32B model at Q4_K_M. Multiply 32 billion by 4.9 bits and divide by eight: about 19.6 GB, or roughly 18 GiB. That's an arithmetic estimate. Different architectures keep different tensors at higher precision, so check the actual file size on the repository page, which is usually listed next to each download.

IQ quants and the older _0 types

Two other families show up in most repositories, and they follow different rules from the K types.

IQ types and the importance matrix

IQ4_XS, IQ3_S and friends use an importance matrix: a file built by running sample text through the model and recording which weights matter most. The quantizer then spends its bits where they count. The Hugging Face table puts IQ4_XS at 4.25 bits per weight, a little under Q4_K. The quantize README says the accuracy loss from quantizing "can be minimized by using a suitable imatrix file", which is why the smallest usable files at 2 and 3 bits are nearly always IQ types.

The catch is that the quality of an IQ file depends on the sample text used to build its matrix. Two IQ4_XS uploads of the same model from different people can behave differently, so note who made the file.

Q8_0 and Q4_0

The _0 and _1 types are the older, simpler formats: one scale per block of 32 weights and nothing else. Hugging Face marks them as legacy. Q8_0 is still everywhere, though, because at eight bits there's so little loss that the simple method is good enough. In the Llama 3.1 8B table it's 8.5 bits per weight and 7.95 GiB. Q4_0 is a different story. At four bits the K and IQ types do better for the same size, so you rarely need it unless your runtime or hardware only supports the old types.

Picking a quant for your card

Knowing the names makes the choice shorter. You still have to try the file on your own work.

Start from memory

Work out what you can afford first. The VRAM calculator takes your GPU and a model and shows the weights, the context and the runtime overhead as separate lines. Pick the largest quant that leaves room for the context you really use. A Q5_K_M that fits with 4K of context is a worse choice than a Q4_K_M that fits with the 16K your documents need. Our VRAM requirements guide covers where the extra memory goes.

Then test quality on your own work

Bit width tells you how much was thrown away. It can't tell you whether your tasks notice. Run the same prompts on two neighbouring quants, say Q4_K_M and Q5_K_M, and compare the answers you'd accept. If they tie, keep the smaller one and enjoy the headroom. The quantization guide explains the trade-offs in more depth, and the GPU reports show which quants people run on cards like yours. When you settle on one, share your setup with the exact file name, so the next person with your card doesn't have to guess.

Frequently asked questions

What does Q4_K_M mean?

Q4 is the base bit width, about four bits for most weights. K means llama.cpp's k-quant method with block scales. M is the medium recipe, which stores a few sensitive tensors at 6 bits. On Llama 3.1 8B it averages about 4.9 bits per weight.

What's the difference between Q4_K_S and Q4_K_M?

Q4_K_S keeps every tensor at the 4-bit type. Q4_K_M moves part of the attention value and feed-forward weights up to 6 bits. On Llama 3.1 8B that costs about 0.2 GiB more. Most people pick M unless those 200 MB decide whether the model fits.

Is a smaller model at Q8_0 smarter than a bigger one at Q4_K_M?

There's no general rule. A bigger model at 4 bits often knows more, and a smaller one at 8 bits loses almost nothing to quantization. Run both on the tasks you care about and keep whichever passes more of them at a speed you can live with.

How much VRAM does a Q4_K_M model need?

Start with the file size, then add the context cache and the runtime's own buffers. A 4.6 GiB file won't run comfortably on a 6 GB card once you add a long context. The llamaperf calculator adds these up for your card and context length.

What are IQ quants?

IQ types such as IQ4_XS and IQ3_S use an importance matrix built from sample text to decide which weights keep more precision. They give better quality than K types at the same very small sizes, and their quality depends on the sample text the uploader used.

Why is my Q4_K_M file bigger than 4 bits per weight suggests?

Each block of weights also stores scales, and the M recipe keeps some tensors at 6 bits. Together that pushes Q4_K_M to about 4.9 bits per weight. Some models also keep embeddings or output layers at higher precision, which adds more.