Context Length
What This Article Covers
- What tokens and context windows are, and how they relate
- How much text actually fits in a context window, with concrete numbers
- Why the KV-cache consumes so much memory with long contexts
- What context lengths common local models offer
- What happens when your input exceeds the context window
Introduction: Context Length Explained
Language models don’t process text as whole words, but as tokens. The number of tokens a model can handle in a single pass is called the context length or context window. This value determines how much information you can give the model at once. For RAG, long documents, code files, or extended chat histories, it’s one of the most critical parameters.
Anyone running local AI runs into this limit immediately. A model with 4,000 token context can’t read a long contract in one go. A model with 128,000 tokens handles it easily, but consumes significantly more memory in the process. This article explains what’s behind it, how to weigh the right values, and what to consider for your hardware.
Why Do You Need to Understand Context Length?
Imagine you want to summarize a 20-page contract. You load the document into your local model, write your prompt, and hit go. Instead of a summary, you either get an error message, an incomplete response, or the model starts midway through the text and completely ignores the beginning.
What went wrong? The contract might span 8,000 tokens. But your model only has a context window of 4,000 tokens. It simply can’t capture the entire text at once. The model either takes the first 4,000 tokens and cuts off the rest, or it fails with an error. Either way, the result is useless.
This scenario isn’t theoretical. It happens daily once you work with real documents. Without understanding context length, you’re left puzzled about why the model seemingly ignores text at random or cuts off mid-sentence. With this knowledge, you spot the problem immediately and can fix it, whether through chunking, summarization, or choosing a model with a larger context window.
Context Length at a Glance
Context length is the maximum number of tokens a model can process in a single step. Think of the context window like the size of your desk. You can only spread out as many documents simultaneously as space allows. Anything that doesn’t fit on the desk must be set aside or summarized before you continue.
A larger context window means more desk space. You can submit longer documents, entire code files, or longer conversation histories at once. The catch: the longer the context, the more memory the model needs during inference. The KV-cache in particular grows substantially and quickly becomes a bottleneck on consumer hardware.
Who Is This For?
This article targets beginners experimenting with local AI and wanting to understand why some models handle long text while others don’t. If you’re using Ollama, LM Studio, or similar tools and wondering why your model struggles with long documents, you’re in the right place. No background in math or computer science required, just a basic grasp of how AI processes text.
Key Terms Around Context Length
| Term | Definition |
|---|---|
| Token | Small text unit processed by the model, usually a word part, whole word, or character |
| Context Window | Maximum number of tokens a model can accept in a single step |
| KV-Cache | Buffer for already-processed tokens, speeds up generation, grows with context length |
| Prompt | The input you give the model, consisting of system prompt and user messages |
| Completion | The response generated by the model |
| System Prompt | An initial instruction that controls the model’s behavior |
| Context Window | Maximum token capacity for a single inference pass |
| Tokenizer | The tool that breaks text into tokens, model-specific |
What Is a Token?
A token is the smallest unit a language model processes. One token is not one word. Often a token represents part of a word, sometimes just a single character. The tokenizer, a model-specific component, breaks your text into these units.
Example of how an English sentence is split:
"Hello world, how are you?"
The tokenizer might produce these tokens:
["Hel", "lo", " world", ",", " how", " are", " you", "?"]
Every model comes with its own tokenizer. A token in a Llama model doesn’t have to match a token in a Qwen model. This means the same text produces different token counts across models.
This becomes especially relevant for German. English models typically train on English text, so English words often encode as single tokens. German words, less common in training data, get split across multiple tokens.
A quick comparison:
| Language | Text | Approximate Token Count |
|---|---|---|
| English | ”The quick brown fox jumps over the lazy dog.” | roughly 10 tokens |
| German | ”Der schnelle braune Fuchs springt über den faulen Hund.” | roughly 14 tokens |
For German, plan on roughly 1.3 to 1.5 tokens per word. A 1,000-word text corresponds to about 1,300 to 1,500 tokens. This matters when you estimate whether a document fits in a context window.
How Does the Context Window Work?
The context window is the entire space available for input and output. Everything must fit inside: the system prompt, all previous user messages, all model responses, and the new completion being generated.
The formula is straightforward:
Context Window = System Prompt + User Messages + Model Responses + New Completion
If your model has an 8,000-token context window and your system prompt already uses 500 tokens, you have 7,500 tokens left for the actual dialogue. If you then submit a document with 6,000 tokens, the model has only 1,500 tokens remaining for its response. With longer chat histories, space fills up quickly.
This means a large context window matters not just for input, but for output too. If the model needs to write a long response, it requires free tokens in the window. If input is too long, little space remains for the answer, and the model cuts off mid-sentence.
The KV-Cache Explained
The KV-cache is one of the most important yet frequently misunderstood aspects of context length. KV stands for Key and Value, two types of intermediate values the model calculates and stores for each token during processing.
Why does this cache exist? During text generation, the model calculates attention values across all previous tokens for each new one. Without the cache, it would have to recalculate all prior tokens for every new token. That would be extremely slow. The KV-cache stores the computed values instead, so the model can reuse them.
The problem: the cache grows with every token. For each token, the model stores values in every layer and at every attention head. KV-cache size depends on three factors:
- Number of tokens in context
- Number of model layers
- Size of attention heads
A concrete example: with a model having 32 layers and a 32,000-token context, the KV-cache can reach several gigabytes. At 128,000 tokens, it grows further. This is why long contexts on consumer hardware often work only with smaller or quantized models.
Quantization reduces model weight size, but the KV-cache still grows with sequence length. Approaches like KV-cache quantization compress the cache itself, but that’s an advanced topic. For now, remember this: the longer the context, the more memory required during inference.
Context Lengths at a Glance
Different models offer vastly different context windows. Here’s an overview of popular models suitable for local deployment:
| Model | Context Window | Notes |
|---|---|---|
| Llama 3.1 (8B, 70B) | 128,000 tokens | Very large window, excellent for lengthy documents |
| Llama 3 (8B) | 8,192 tokens | Standard window, sufficient for most tasks |
| Qwen 2.5 (7B, 14B, 32B) | 128,000 tokens | Large window, strong German language support |
| Mistral 7B / v0.3 | 32,000 tokens | Mid-sized window, good balance |
| Mixtral 8x7B | 32,000 tokens | Mixture-of-Experts, higher memory requirements |
| Phi-3 Mini | 4,096 tokens | Compact and fast, but limited window |
| Gemma 2 (9B, 27B) | 8,192 tokens | Solid window for the model size |
These numbers represent the maximum theoretical context length. In practice, available VRAM and inference speed typically determine how much context you can realistically use. A 128K window won’t help if your VRAM fills up after 16,000 tokens.
What Happens When Input Exceeds the Context Window?
When your input surpasses the context limit, you’ll typically encounter one of three behaviors:
1. Truncation: Many tools automatically cut off the text, keeping the first tokens and discarding the rest. The model only sees part of your document, and you may not even realize it’s happened. Your answer then rests on incomplete information.
2. Error Message: Some tools and models stop with an error. This is honest but not always practical. You’ll need to shorten or split the text yourself.
3. Automatic Chunking: Advanced tools like RAG pipelines automatically divide text into chunks and process them sequentially. This works, but it adds complexity and can lose information when meaning spans chunk boundaries.
Ollama often throws an error on overly long input. Other tools like LM Studio silently truncate. It’s worth understanding how your tool behaves before working with large documents.
Context Length and RAG
RAG stands for Retrieval-Augmented Generation, a technique where the model retrieves relevant text passages from a database before generating its answer. Context length plays a central role here.
In RAG, you search a vector database for matching text chunks and pass them to the model alongside your question. More chunks mean more context needed. A larger context window lets you send more and longer chunks, which can improve answer quality.
But beware: more context isn’t automatically better. Models often struggle with the “Lost in the Middle” problem, where they miss relevant information buried deep in a long context. Memory usage also grows, and inference time increases.
For local RAG: choose a model with adequate but not excessive context. For most RAG applications, 8,000 to 32,000 tokens suffice. If you need more, explore Local RAG and specifically Chunking to prepare your texts optimally.
Strategies for Working Within Limited Context
If your model has a small window or limited VRAM, several approaches help:
Chunking: Split long documents into smaller pieces. Process each chunk individually and synthesize the results. This is the standard RAG method and works without a vector database too. Learn more in Chunking.
Progressive Summarization: Before feeding a long text to the model, have it summarize in stages. Summarize the first half, then the second half, then combine both summaries before passing them to the model.
Smart Prompting: Think carefully about what information the model actually needs. Instead of uploading an entire book, provide only relevant chapters. A precise prompt with the right details often works better than dumping raw text.
Trim Chat History: In long conversations, messages accumulate. Delete old or irrelevant messages, or summarize prior conversation before continuing. This saves context and often improves answer quality too.
Optimize System Prompt: A lengthy system prompt permanently occupies space in the context window. Keep it as concise as possible without losing essential instructions.
Common Context Length Pitfalls
- Conflating tokens and words: One word does not equal one token. German typically needs 1.3 to 1.5 tokens per word. Calculating in words overestimates the model’s capacity.
- Forgetting the system prompt: The system prompt takes up context space. If you don’t account for it, you’ll wonder why the model cuts off sooner than expected.
- Ignoring the KV cache: A large context window is useless if VRAM runs out. The KV cache grows with sequence length and can exhaust memory before hitting the theoretical limit.
- Assuming more context is always better: Long contexts demand more time and memory. Moreover, models sometimes struggle to retain relevant information in very long inputs. For many tasks, 8,000 tokens is sufficient.
- Missing silent truncation: Some tools truncate long input without warning. The model appears to answer correctly but works from incomplete data. Check if your tool provides warnings.
- Letting chat history grow unchecked: In extended conversations, the history fills the context window until little room remains for answers. Regular pruning or summarization is essential.
- Wrong expectations from quantization: Quantization reduces model weights but doesn’t significantly shrink the KV cache, which depends on sequence length and architecture. Loading a quantized 8B model with 128K context still demands substantial memory for the cache alone.
Hardware, Costs, and Security with Long Contexts
Long contexts require more VRAM or RAM. If you process many lengthy documents, you’ll either need large models and powerful workstations or deliberately reduce inputs. Details are in RAM and VRAM Requirements.
Quantization helps with model weights but doesn’t significantly reduce KV cache size, since cache depends on sequence length and model architecture. If you load an 8B model with 128K context, you still need substantial memory just for the cache.
In local AI, costs mainly come from hardware and electricity. The advantage: even long documents never leave your network. This matters for confidential contracts, internal code, or personal data. Cloud services charge per token, making long contexts expensive.
Local deployment has a clear security advantage. You don’t send sensitive data to external servers. In return, you’re fully responsible for backups, updates, and access control. See Running LLMs Locally for more.
Further Reading and Resources on Context Length
- What is local AI?
- Quantization
- RAM and VRAM Requirements
- Running LLMs Locally
- Local RAG
- Chunking
- Ollama
FAQ - Common Questions About Context Length
How many words equal one token?
Roughly 0.75 words per token in English, or 0.6 to 0.7 words per token in German. A German text with 1,000 words typically translates to about 1,300 to 1,500 tokens.
Can I input more than the context window allows?
No. The model will either ignore tokens beyond the window or produce errors. Tools truncate or split the text beforehand, sometimes without warning.
Why does the KV cache grow so large?
It stores intermediate values for every processed token across each model layer. With longer texts, these values accumulate and consume significant memory, scaling linearly with sequence length.
Are longer context windows always better?
Not necessarily. Larger contexts consume more memory and processing time. Additionally, models sometimes overlook relevant information in the middle of very long inputs. For many tasks, 8,000 to 32,000 tokens suffice.
What happens to old chat messages?
When the conversation history exceeds the context window, older messages get truncated. Many applications work around this by summarizing the conversation so far to save space.
How do I find out how many tokens my text contains?
Use an online token counter or the tokenizer libraries from Hugging Face. Make sure to use the tokenizer for your specific model, since different models tokenize text differently.
Does context length affect speed?
Yes. Longer contexts require more computation per token because the model calculates attention across all previous tokens. Response time increases with context length.
Can I quantize the KV cache?
Yes, approaches like KV cache quantization can compress the cache. However, this is an advanced technique and not available in every tool. For now, it’s enough to know that the cache grows and consumes memory.
What context window do I need for RAG?
For most RAG applications, 8,000 to 32,000 tokens work well. You input a few text chunks plus your question, and that’s usually sufficient. More context only helps if you want to pass many or very long chunks.
Why does my model cut off mid-response?
The context window is likely full. Input and output share the same window. If the input is too long, there’s no room for the response and the model stops. Either shorten the input or switch to a model with a larger context window.
Does quantization help with long contexts?
Quantization reduces model weights, but the KV cache still grows with sequence length. You save memory on the weights, but with very long contexts, the cache can still become a bottleneck.
What does “Lost in the Middle” mean?
This is an observed effect where models pay less attention to information in the middle of a long context compared to the beginning or end. If you have important information, place it near the start or end of your prompt.
Can I increase a model’s context window?
No, the context window is fixed by training. A model trained on 8,000 tokens can’t suddenly handle 32,000 tokens. You need to choose a model that supports the desired window size from the start.
Do I need special hardware for 128K context?
Not special hardware per se, but you need enough. The KV cache for 128,000 tokens can occupy several gigabytes. You need sufficient VRAM or RAM to hold both the model and cache simultaneously. Running locally typically requires 16 GB of VRAM or more.
Sources and Further Reading
- Hugging Face Tokenizer Documentation
- Ollama Model Overview
- Llama 3.1 Context Length Specification
- Qwen 2.5 Model Card
- Mistral AI Documentation
- Research on “Lost in the Middle” (Liu et al., 2023)


