Skip to content
BotServBotServ
OllamaPerformanceGPUQuantizationOptimization

Optimize Ollama Performance

Speed up Ollama with GPU optimization, context length, quantization, batch sizes, and tuning tips.

S

schutzgeist

3 min read
Optimize Ollama Performance

Optimizing Ollama Performance

What this article covers

  • Identifying where Ollama becomes a bottleneck.
  • Checking GPU utilization and driver status.
  • How quantization and model size affect speed.
  • Tuning context length, batch size, and parameters.
  • Key optimization tips and common pitfalls.

Introduction: Optimizing Ollama Performance

Ollama is straightforward to use, but speed depends heavily on hardware, model size, and configuration. Anyone running larger models or processing long documents quickly discovers that response times climb. A few targeted adjustments can often significantly speed up Ollama without requiring new hardware.

This article walks through where to focus your Ollama optimization efforts and which levers deliver the biggest gains.

Key terminology

  • Prompt-Eval: Processing the input through the model.
  • Generation: Producing the response.
  • Tokens per second: Metric for generation speed.
  • Context Window: Maximum number of tokens in context.
  • KV-Cache: Cache for already-processed tokens.
  • Batch Size: Number of tokens processed in parallel.
  • Quantization: Reducing the bit depth of model weights.
  • Offload: Distributing computation between CPU and GPU.

Causes of poor performance

  • Model too large: Doesn’t fit in VRAM, runs partially on CPU.
  • Overly long contexts: More context means more computation.
  • CPU instead of GPU: Ollama doesn’t detect GPU due to missing drivers.
  • Slow memory: System RAM too slow or not in dual-channel mode.
  • Wrong model variant: Running a 70B model when an 8B version exists.
  • Too many models loaded: VRAM becomes fragmented.

Checking GPU utilization

ollama ps
nvidia-smi

If Ollama isn’t running on your GPU, verify your drivers. For Nvidia, the proprietary driver must be installed. For AMD, ROCm is required.

Installing drivers

Nvidia

sudo apt install -y nvidia-driver-535 nvidia-utils-535

AMD

sudo apt install -y amdgpu-install
sudo amdgpu-install -y --usecase=rocm

Note: Older AMD GPUs like GCN 1.0 are not supported by ROCm.

Choosing model size

For most tasks, smaller models are sufficient:

  • Qwen 2.5 7B or 14B: Strong general-purpose performance.
  • Llama 3.1 8B: Broad compatibility.
  • Mistral 7B or Nemo 12B: Solid multilingual support.

Larger models are better quality, but only if they fit entirely in VRAM. Otherwise, you’ll see worse performance than with smaller models that run fully on GPU.

Quantization

4-bit quantization is the standard. With limited VRAM, try 3-bit or 2-bit, though quality suffers. For coding and reasoning tasks, 5-bit often performs better.

ollama pull qwen2.5:14b-q4_K_M
ollama pull qwen2.5:14b-q5_K_M

Setting context length

Not every model supports every context size. Overly large num_ctx values slow processing. Often 2048 or 4096 tokens is enough.

PARAMETER num_ctx 4096

In the Modelfile:

FROM qwen2.5:14b
PARAMETER num_ctx 4096

Ollama environment variables

In /etc/systemd/system/ollama.service.d/override.conf or in your Docker setup:

[Service]
Environment="OLLAMA_NUM_PARALLEL=4"
Environment="OLLAMA_MAX_LOADED_MODELS=2"
Environment="OLLAMA_GPU_OVERHEAD=1GB"

Then reload:

sudo systemctl daemon-reload
sudo systemctl restart ollama

Parallelism

OLLAMA_NUM_PARALLEL controls how many requests are handled concurrently. Useful for server deployments; for single-user setups, typically set to 1.

Accelerating CPU inference

When no GPU is available:

  • Use a recent CPU with AVX2 or AVX512 support.
  • Allocate sufficient RAM.
  • Allow multiple threads:
PARAMETER num_thread 8

Monitoring

Check regularly:

ollama ps
free -h
nvidia-smi

If the model spills to disk, performance degrades.

Common pitfalls

  • GPU not detected: Missing or incorrect driver version.
  • Model runs on CPU: Too large for available VRAM.
  • Context length too high: Slows processing and causes errors.
  • Multiple models loaded simultaneously: Fragmented VRAM.
  • Outdated Ollama version: Updates often bring performance improvements.

Further reading and resources

FAQ: Ollama Performance

How do I check if Ollama is using my GPU? ollama ps shows whether a model uses CPU or GPU.

Why is my GPU slow? The model may be too large and running partially on CPU.

Does more VRAM help? Yes, especially to fit larger models entirely on GPU.

Is 3-bit quantization recommended? Only if 4-bit doesn’t fit in memory, since quality suffers.

Should I set num_ctx to the maximum? No, only use as much context as actually needed.

Sources and further reading

Summary: Optimizing Ollama Performance

Ollama performance improves significantly through the right model size, quantization, drivers, context length, and Ollama parameters. The golden rule: your model should fit entirely in GPU memory. Bigger isn’t always better if it forces computation to the CPU. Monitor ollama ps, nvidia-smi, and system resources regularly, and you’ll quickly spot the largest bottlenecks.

Back to Blog
Share:

Related Posts