llamaperf
← All articles

Hardware planning

Why a longer context needs more memory

By llamaperf · · 5 min read

Quick answer

The model keeps keys and values for every token in the context, in every layer, so the cache grows in step with the context you set. Size the context to your longest routine input. If memory is tight, an 8-bit KV cache in Ollama takes about half the space of the default 16-bit one.

What the KV cache holds

A language model writes one token at a time, and each new token has to look back at every token before it. Recomputing all of that on every step would be painfully slow, so runtimes keep the intermediate results around.

A KV cache stores these calculations so they can be reused without recomputing them. From the Hugging Face Transformers cache documentation

One entry per token, per layer

For every token in the context, each attention layer stores two vectors, the keys and the values. That's the KV in KV cache. A model with 32 layers keeps 32 pairs for every token you've sent and every token it has written. The weights stay the same size whatever you do. The cache is the part that grows while you talk to it.

Why it grows in a straight line

Here's the arithmetic for a model with 32 layers, 8 key-value heads and a head size of 128, stored at 16 bits (2 bytes) per value:

2 (key and value) x 32 layers x 8 heads x 128 x 2 bytes = 131,072 bytes per token
131,072 bytes x 32,768 tokens = 4 GiB

Double the context and you double that 4 GiB. These are example numbers, and your model's shape is in its config file. The Hugging Face page puts it plainly: the cache "can occupy a significant portion of memory and become a bottleneck for long-context generation."

The window you set and the text you send

Two numbers get mixed up here. One is the context window you configure. The other is how many tokens you send.

Runtimes reserve the window up front

llama.cpp sizes its cache for the context you pass with -c when the server starts, so a 32K setting costs 32K worth of memory even if you only ever ask short questions. Ollama's FAQ says it uses a 4096-token window by default, and you raise it with OLLAMA_CONTEXT_LENGTH or the num_ctx parameter. That default is why a long document can get quietly cut off in Ollama until you change it.

The same FAQ warns that parallel requests multiply the cost. It says required RAM scales with OLLAMA_NUM_PARALLEL times OLLAMA_CONTEXT_LENGTH, so four parallel slots at 16K need the cache of a 64K window.

Some models stop growing

Not every layer keeps the whole history. Models with sliding-window attention only keep the last few thousand tokens in those layers, and the Hugging Face docs note that for them "the cache will stop growing when the layers using these types of attention have reached their maximum size". That's why two models of the same size can need very different amounts of memory at 32K.

Making the cache smaller

If the model fits but the context you need doesn't, you have two settings to try before you reach for a smaller model.

Quantize the KV cache

Ollama can store the cache at lower precision through the OLLAMA_KV_CACHE_TYPE environment variable. Its FAQ gives the trade: q8_0 uses about half the memory of the default f16 with a very small loss in precision, and q4_0 uses about a quarter with a loss that may show up more at long contexts. The setting is global, so every model you load picks it up. In llama.cpp the same idea is -ctk and -ctv, the cache types for keys and values.

Start with q8_0. Then rerun the task that depends most on details from early in a long input, because that's where a lower-precision cache goes wrong first.

Turn on Flash Attention

Ollama only quantizes the cache when Flash Attention is on, and the FAQ says Flash Attention itself "can significantly reduce memory usage as the context size grows". Recent versions switch it on automatically when your hardware supports it. OLLAMA_FLASH_ATTENTION=1 forces it.

Size the context to your work

The cheapest memory saving is a context that matches what you do.

Measure your longest input

Take the longest thing you paste in a normal week, say a contract or a long chat, and count its tokens with your runtime or tokenizer. Add room for the answer and a margin. For quick questions 8K is plenty. For document work you might need 32K. Setting 128K because the model supports it just spends memory you could have used for a better quant.

Check the fit before you commit

Put your card, the model and that context into the VRAM calculator. It shows the weights and the cache on separate lines, so you can see which one is squeezing you. If it's the weights, our post on reading GGUF quant names helps you pick a smaller file. On a Mac, remember the model and the cache share one memory pool with everything else you run, so check the Mac reports from people with the same memory size. The VRAM requirements guide covers the rest of the budget.

Frequently asked questions

How do I find out how much VRAM the context needs?

Multiply two by the layer count, the key-value head count, the head size and the bytes per value. That gives the bytes per token. Multiply by your context length. The llamaperf calculator does this for you and shows the cache next to the weights.

How much can I quantize my KV cache before quality drops?

Ollama's documentation describes q8_0 as a very small loss at half the memory, and q4_0 as a small to medium loss that shows more at long contexts. Try q8_0 first and test the task that relies on details from early in a long input.

Does the context length setting matter if I only ask short questions?

Yes for memory. llama.cpp reserves the cache for the whole window when it starts, so a big setting costs memory even when your prompts are short. It doesn't change the answers to short questions.

What's a good context length for a personal assistant?

Size it to your longest routine input plus room for the reply. Short chats rarely need more than 8K. If you paste documents or long code files, measure the biggest one and set the window just above it.

Why does my model use more memory than the file size?

The file holds the weights. At runtime the model also needs the KV cache for your context and compute buffers. At long contexts the cache alone can take several gigabytes.