Understanding Context Length in Ollama
What This Article Covers
- What context length means.
- How tokens are counted.
- How
num_ctxworks in Ollama. - The relationship between context, VRAM, and speed.
- Choosing the right length for your use case.
Introduction: Understanding Context Length in Ollama
Context length describes how many tokens a model can process at once. This includes the system prompt, your question, chat history, and RAG documents. A larger context lets the model draw on more background information. However, longer context requires more VRAM and compute time. Setting context length too high risks slowdowns or crashes.
This article explains how Ollama handles context length and how to configure it effectively.
Key Terms
- Token: A small text unit, roughly a word fragment.
- Context length: The number of tokens a model processes.
- num_ctx: The Ollama parameter for context length.
- VRAM: GPU video memory.
- Prompt-Eval: Processing the input context.
- KV-Cache: Buffer for attention values.
- Truncation: Cutting off overly long contexts.
- Sliding Window: A bounded window for older tokens.
Counting Tokens
Tokens are not the same as words. A single word can split into multiple tokens. Short words often map to one token, while longer or compound words take several. German umlauts and special characters can create additional tokens.
Examples:
- “Hallo” = 1-2 tokens.
- “BotServ.de” = 3-4 tokens.
- “Künstliche Intelligenz” = roughly 4-6 tokens.
num_ctx in Ollama
num_ctx sets the maximum context length:
ollama run llama3.1
>>> /set parameter num_ctx 8192
Or via Modelfile:
PARAMETER num_ctx 8192
Or via API:
{
"model": "llama3.1",
"prompt": "...",
"options": {
"num_ctx": 8192
}
}
Default Context
Many models come with a default value, often 2048 or 4096. This conserves VRAM but limits context. In Ollama, num_ctx can often be set well above the default if your model and hardware support it.
VRAM Requirements
Context length and VRAM are linked:
VRAM ~ Model size + KV-Cache
KV-Cache grows with num_ctx * (layers + hidden_size)
A context twice as long requires significantly more memory. With 7B models and 4096 tokens, 8-12 GB often suffices. At 32k tokens, you may need 20-30 GB.
Speed
A long context primarily slows down prompt evaluation. Generation itself can also become slower because more KV-Cache per layer must be read. For real-time chat, context should not be unnecessarily large.
When You Need Large Context
- Summarizing long documents.
- Chat with extensive history.
- RAG with many chunk overlaps.
- Code reviews spanning multiple files.
When Small Context Is Enough
- Short question-answer scenarios.
- Fast chatbots.
- Limited memory.
- Quick responses matter more than background.
Practical Values
| Use Case | num_ctx |
|---|---|
| Short chat | 2048 |
| Long chat | 4096-8192 |
| Document RAG | 8192-16384 |
| Extended code review | 16384-32768 |
| Very long documents | 32768+ |
Not every model supports every length. Check the model documentation.
Context Too Large
Consequences:
- Slow responses.
- Out-of-memory errors, model crashes.
- Error messages from insufficient VRAM.
- Poor quality, as the model forgets older sections.
Limiting Context
- Include only relevant documents in the prompt.
- Reduce chunk size.
- Summarize chat history.
- Set
num_ctxrealistically.
Counting and Testing
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "...long text...",
"options": {
"num_ctx": 8192
},
"stream": false
}'
Monitor VRAM during the call:
watch -n 1 nvidia-smi
Tips
- Keep
num_ctxno larger than necessary. - Respect model limits.
- Estimate tokens beforehand.
- Split long documents into chunks.
- Summarize chat history regularly.
- Benchmark with different lengths.
Further Reading and Resources
- BotServ.de Ollama Performance
- BotServ.de Ollama Model Benchmarks
- BotServ.de Ollama RAG
- BotServ.de Ollama Modelfiles
FAQ: Context Length in Ollama
What is num_ctx? An Ollama parameter that sets the maximum context length.
How many tokens fit in 8 GB of VRAM? With a 7B model at Q4 quantization, typically 2k-8k depending on the model.
Can I increase num_ctx arbitrarily? Only up to the limit supported by your model and hardware.
Why does the model slow down with long context? Because more KV-Cache must be computed and read.
How many tokens does my prompt have?
Check with ollama show or external tokenizer tools.
Sources and Further Reading
- Ollama Modelfile: https://github.com/ollama/ollama/blob/main/docs/modelfile.md
- Tokenizer Visualization: https://platform.openai.com/tokenizer
- KV Cache: https://arxiv.org/abs/1911.02150
Context length determines how much background information Ollama can consider. Larger context helps with long documents or chat history, but costs VRAM and processing time. The num_ctx parameter provides straightforward control. When you balance context, model size, and hardware requirements, you get good results without unexpected performance drops.


