Optimizing Ollama Performance
What this article covers
- Factors that slow down Ollama.
- Flash Attention and ggml optimizations.
- KV-Cache and its impact.
- Quantization and model size.
- Practical configuration and commands.
Introduction: Optimizing Ollama Performance
Ollama is straightforward to use, but getting maximum speed from your hardware requires understanding a few key details. Slow token throughput often stems from suboptimal quantization, unnecessarily large context windows, CPU rather than GPU execution, or missing optimizations like Flash Attention. With the right configuration, you can achieve significant improvements in most cases.
This article walks through the levers you can pull to make Ollama faster.
Key terms
- Flash Attention: Memory-efficient attention computation.
- KV-Cache: Buffer for key-value pairs.
- Quantization: Reducing model precision.
- num_gpu: Number of model layers to offload to GPU.
- num_thread: CPU threads.
- batch_size: Number of tokens processed simultaneously.
- mlock: RAM reservation to prevent swapping.
- f16_kv: KV-Cache in half-float format.
Main optimization levers
Use GPU instead of CPU
ollama run llama3.1
If Ollama runs on CPU, verify your setup with:
nvidia-smi
or
rocm-smi
Ollama must detect CUDA or ROCm support.
Force more GPU layers
ollama run llama3.1 --verbose
Via API call:
{
"options": {
"num_gpu": 50
}
}
num_gpu controls how many model layers are loaded onto the GPU.
Enable Flash Attention
Ollama uses Flash Attention when supported. Some models may require special compilation or a ggml build with Flash Attention enabled. Newer Ollama versions often activate it automatically.
Reduce quantization
Lower quantization uses less VRAM and computes faster:
- Q4_0: fast, lower quality.
- Q4_K_M: good balance.
- Q5_K_M: better quality.
- Q8_0: large file, high quality.
ollama run llama3.1:8b-instruct-q4_K_M
Reduce context window
PARAMETER num_ctx 4096
More context equals slower performance. Keep it only as large as you need.
Limit num_predict
PARAMETER num_predict 512
This prevents unnecessarily long responses and bounds computation time.
Adjust CPU threads
For pure CPU operation:
ollama run llama3.1
Or via Modelfile:
PARAMETER num_thread 8
Use mlock
Prevents Ollama from being swapped to disk:
PARAMETER use_mlock true
Monitor memory usage
Watch resource consumption during execution:
watch -n 1 nvidia-smi
If VRAM is maxed out, reduce model size, quantization, or context length.
Batch size
Longer inputs process faster when num_batch is increased:
PARAMETER num_batch 512
Not all models and backends benefit equally, so test your specific setup.
Set temperature to 0
For reproducible and slightly faster test runs:
PARAMETER temperature 0
Update Ollama
ollama -v
New releases often include performance improvements.
Example Modelfile
FROM llama3.1:8b-instruct-q4_K_M
PARAMETER num_ctx 4096
PARAMETER num_gpu 50
PARAMETER num_predict 512
PARAMETER use_mlock true
PARAMETER num_batch 512
PARAMETER temperature 0.7
Tips
- Always check GPU utilization first.
- Choose the smallest quantization that meets your quality needs.
- Don’t oversizethe context window.
- Keep Ollama updated.
- Benchmark before and after optimizations.
- Monitor hardware temperature.
Common pitfalls
- Running on CPU instead of GPU: Usually a driver or Ollama installation issue.
- Context window too large: Fills memory, performance degrades.
- Unsupported quantization: Model falls back to CPU.
- Excessive KV-Cache: Slower prompt evaluation.
- Outdated Ollama version: Missing Flash Attention.
- Thermal throttling: Performance drops after extended load.
Further reading
- BotServ.de Ollama Model Benchmarks
- BotServ.de Ollama Performance
- BotServ.de Ollama Context Length
- BotServ.de Ollama Quantization in Practice
FAQ: Ollama Performance
Why is Ollama slow? Usually CPU execution, model too large, high context, or thermal throttling.
What’s the best way to make Ollama faster? Use GPU, choose lower quantization, reduce context, keep Ollama up to date.
Does Flash Attention make a big difference? Yes, especially with long contexts, though not always available.
Should I adjust num_gpu? It can help if automatic layer distribution isn’t optimal.
Does more CPU RAM help? Only if you’re computing on CPU. VRAM matters more for GPU execution.
Sources and further reading
- Ollama Performance: https://github.com/ollama/ollama/blob/main/docs/troubleshooting.md
- Flash Attention: https://github.com/Dao-AILab/flash-attention
- llama.cpp: https://github.com/ggerganov/llama.cpp
Summary: Optimizing Ollama Performance
Ollama’s speed improves significantly through GPU acceleration, appropriate quantization, context window reduction, and cache optimization. Flash Attention and staying current with Ollama versions provide additional gains. Regularly benchmarking and tuning parameters like num_ctx, num_predict, and num_batch helps you find the optimal balance for your hardware and use case.


