Skip to content
BotServBotServ
OllamaSpeedPerformanceFlash AttentionKV-Cache

Optimize Ollama Speed

Speed up Ollama with Flash Attention, KV-Cache, quantization, GPU optimization and parameter tuning.

S

schutzgeist

3 min read
Optimize Ollama Speed

Optimizing Ollama Performance

What this article covers

  • Factors that slow down Ollama.
  • Flash Attention and ggml optimizations.
  • KV-Cache and its impact.
  • Quantization and model size.
  • Practical configuration and commands.

Introduction: Optimizing Ollama Performance

Ollama is straightforward to use, but getting maximum speed from your hardware requires understanding a few key details. Slow token throughput often stems from suboptimal quantization, unnecessarily large context windows, CPU rather than GPU execution, or missing optimizations like Flash Attention. With the right configuration, you can achieve significant improvements in most cases.

This article walks through the levers you can pull to make Ollama faster.

Key terms

  • Flash Attention: Memory-efficient attention computation.
  • KV-Cache: Buffer for key-value pairs.
  • Quantization: Reducing model precision.
  • num_gpu: Number of model layers to offload to GPU.
  • num_thread: CPU threads.
  • batch_size: Number of tokens processed simultaneously.
  • mlock: RAM reservation to prevent swapping.
  • f16_kv: KV-Cache in half-float format.

Main optimization levers

Use GPU instead of CPU

ollama run llama3.1

If Ollama runs on CPU, verify your setup with:

nvidia-smi

or

rocm-smi

Ollama must detect CUDA or ROCm support.

Force more GPU layers

ollama run llama3.1 --verbose

Via API call:

{
  "options": {
    "num_gpu": 50
  }
}

num_gpu controls how many model layers are loaded onto the GPU.

Enable Flash Attention

Ollama uses Flash Attention when supported. Some models may require special compilation or a ggml build with Flash Attention enabled. Newer Ollama versions often activate it automatically.

Reduce quantization

Lower quantization uses less VRAM and computes faster:

  • Q4_0: fast, lower quality.
  • Q4_K_M: good balance.
  • Q5_K_M: better quality.
  • Q8_0: large file, high quality.
ollama run llama3.1:8b-instruct-q4_K_M

Reduce context window

PARAMETER num_ctx 4096

More context equals slower performance. Keep it only as large as you need.

Limit num_predict

PARAMETER num_predict 512

This prevents unnecessarily long responses and bounds computation time.

Adjust CPU threads

For pure CPU operation:

ollama run llama3.1

Or via Modelfile:

PARAMETER num_thread 8

Use mlock

Prevents Ollama from being swapped to disk:

PARAMETER use_mlock true

Monitor memory usage

Watch resource consumption during execution:

watch -n 1 nvidia-smi

If VRAM is maxed out, reduce model size, quantization, or context length.

Batch size

Longer inputs process faster when num_batch is increased:

PARAMETER num_batch 512

Not all models and backends benefit equally, so test your specific setup.

Set temperature to 0

For reproducible and slightly faster test runs:

PARAMETER temperature 0

Update Ollama

ollama -v

New releases often include performance improvements.

Example Modelfile

FROM llama3.1:8b-instruct-q4_K_M

PARAMETER num_ctx 4096
PARAMETER num_gpu 50
PARAMETER num_predict 512
PARAMETER use_mlock true
PARAMETER num_batch 512
PARAMETER temperature 0.7

Tips

  • Always check GPU utilization first.
  • Choose the smallest quantization that meets your quality needs.
  • Don’t oversizethe context window.
  • Keep Ollama updated.
  • Benchmark before and after optimizations.
  • Monitor hardware temperature.

Common pitfalls

  • Running on CPU instead of GPU: Usually a driver or Ollama installation issue.
  • Context window too large: Fills memory, performance degrades.
  • Unsupported quantization: Model falls back to CPU.
  • Excessive KV-Cache: Slower prompt evaluation.
  • Outdated Ollama version: Missing Flash Attention.
  • Thermal throttling: Performance drops after extended load.

Further reading

FAQ: Ollama Performance

Why is Ollama slow? Usually CPU execution, model too large, high context, or thermal throttling.

What’s the best way to make Ollama faster? Use GPU, choose lower quantization, reduce context, keep Ollama up to date.

Does Flash Attention make a big difference? Yes, especially with long contexts, though not always available.

Should I adjust num_gpu? It can help if automatic layer distribution isn’t optimal.

Does more CPU RAM help? Only if you’re computing on CPU. VRAM matters more for GPU execution.

Sources and further reading

Summary: Optimizing Ollama Performance

Ollama’s speed improves significantly through GPU acceleration, appropriate quantization, context window reduction, and cache optimization. Flash Attention and staying current with Ollama versions provide additional gains. Regularly benchmarking and tuning parameters like num_ctx, num_predict, and num_batch helps you find the optimal balance for your hardware and use case.

Back to Blog
Share:

Related Posts