Skip to content
BotServBotServ
OllamaContextnum_ctxTokenVRAM

Understanding Context Length in Ollama

Context length, token limits, and num_ctx in Ollama. Balance VRAM, speed, and quality.

S

schutzgeist

3 min read
Understanding Context Length in Ollama

Understanding Context Length in Ollama

What This Article Covers

  • What context length means.
  • How tokens are counted.
  • How num_ctx works in Ollama.
  • The relationship between context, VRAM, and speed.
  • Choosing the right length for your use case.

Introduction: Understanding Context Length in Ollama

Context length describes how many tokens a model can process at once. This includes the system prompt, your question, chat history, and RAG documents. A larger context lets the model draw on more background information. However, longer context requires more VRAM and compute time. Setting context length too high risks slowdowns or crashes.

This article explains how Ollama handles context length and how to configure it effectively.

Key Terms

  • Token: A small text unit, roughly a word fragment.
  • Context length: The number of tokens a model processes.
  • num_ctx: The Ollama parameter for context length.
  • VRAM: GPU video memory.
  • Prompt-Eval: Processing the input context.
  • KV-Cache: Buffer for attention values.
  • Truncation: Cutting off overly long contexts.
  • Sliding Window: A bounded window for older tokens.

Counting Tokens

Tokens are not the same as words. A single word can split into multiple tokens. Short words often map to one token, while longer or compound words take several. German umlauts and special characters can create additional tokens.

Examples:

  • “Hallo” = 1-2 tokens.
  • “BotServ.de” = 3-4 tokens.
  • “Künstliche Intelligenz” = roughly 4-6 tokens.

num_ctx in Ollama

num_ctx sets the maximum context length:

ollama run llama3.1
>>> /set parameter num_ctx 8192

Or via Modelfile:

PARAMETER num_ctx 8192

Or via API:

{
  "model": "llama3.1",
  "prompt": "...",
  "options": {
    "num_ctx": 8192
  }
}

Default Context

Many models come with a default value, often 2048 or 4096. This conserves VRAM but limits context. In Ollama, num_ctx can often be set well above the default if your model and hardware support it.

VRAM Requirements

Context length and VRAM are linked:

VRAM ~ Model size + KV-Cache
KV-Cache grows with num_ctx * (layers + hidden_size)

A context twice as long requires significantly more memory. With 7B models and 4096 tokens, 8-12 GB often suffices. At 32k tokens, you may need 20-30 GB.

Speed

A long context primarily slows down prompt evaluation. Generation itself can also become slower because more KV-Cache per layer must be read. For real-time chat, context should not be unnecessarily large.

When You Need Large Context

  • Summarizing long documents.
  • Chat with extensive history.
  • RAG with many chunk overlaps.
  • Code reviews spanning multiple files.

When Small Context Is Enough

  • Short question-answer scenarios.
  • Fast chatbots.
  • Limited memory.
  • Quick responses matter more than background.

Practical Values

Use Casenum_ctx
Short chat2048
Long chat4096-8192
Document RAG8192-16384
Extended code review16384-32768
Very long documents32768+

Not every model supports every length. Check the model documentation.

Context Too Large

Consequences:

  • Slow responses.
  • Out-of-memory errors, model crashes.
  • Error messages from insufficient VRAM.
  • Poor quality, as the model forgets older sections.

Limiting Context

  • Include only relevant documents in the prompt.
  • Reduce chunk size.
  • Summarize chat history.
  • Set num_ctx realistically.

Counting and Testing

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "...long text...",
  "options": {
    "num_ctx": 8192
  },
  "stream": false
}'

Monitor VRAM during the call:

watch -n 1 nvidia-smi

Tips

  • Keep num_ctx no larger than necessary.
  • Respect model limits.
  • Estimate tokens beforehand.
  • Split long documents into chunks.
  • Summarize chat history regularly.
  • Benchmark with different lengths.

Further Reading and Resources

FAQ: Context Length in Ollama

What is num_ctx? An Ollama parameter that sets the maximum context length.

How many tokens fit in 8 GB of VRAM? With a 7B model at Q4 quantization, typically 2k-8k depending on the model.

Can I increase num_ctx arbitrarily? Only up to the limit supported by your model and hardware.

Why does the model slow down with long context? Because more KV-Cache must be computed and read.

How many tokens does my prompt have? Check with ollama show or external tokenizer tools.

Sources and Further Reading

Context length determines how much background information Ollama can consider. Larger context helps with long documents or chat history, but costs VRAM and processing time. The num_ctx parameter provides straightforward control. When you balance context, model size, and hardware requirements, you get good results without unexpected performance drops.

Back to Blog
Share:

Related Posts