Skip to content
BotServBotServ
OllamaFine-tuningLoRAModel adaptationAI training

Fine-tune Ollama Models

Adapt models with Ollama and LoRA techniques. Data prep, training, adapters, and practical examples.

S

schutzgeist

3 min read
Fine-tune Ollama Models

Fine-tuning Ollama Models

What this article covers

  • What fine-tuning is and when it makes sense.
  • How LoRA and QLoRA work.
  • How to prepare training data.
  • How to create an adapter and use it in Ollama.
  • Tips, tools, and common pitfalls.

Introduction: Fine-tuning Ollama models

Off-the-shelf models handle many tasks well, but for specialized domains, particular language styles, or proprietary data, their built-in knowledge often falls short. Fine-tuning adapts a base model to a specific purpose. With Ollama, you can run customized models locally by augmenting them with LoRA adapters. This keeps your data private and lets you steer the model precisely toward your needs.

This article walks through the fundamentals of fine-tuning, data preparation, and integration with Ollama.

Key concepts

  • Fine-tuning: Further training an already-trained model.
  • LoRA: Low-Rank Adaptation, an efficient way to adapt a model with minimal memory overhead.
  • QLoRA: Quantized LoRA, even more memory-efficient.
  • Adapter: Additional weights that augment a base model.
  • Training data: Input-output pairs for learning.
  • Evaluation data: Data used to test the trained model.
  • Overfitting: The model memorizes training data instead of generalizing.

When fine-tuning makes sense

  • You need domain-specific terminology or language patterns.
  • Your own knowledge should inform the model’s responses.
  • You want a particular response format or tone.
  • Simple prompt engineering no longer suffices.
  • RAG helps, but the model needs deeper, more specialized capabilities.

If you have limited data or compute, RAG is often the better first step.

LoRA and QLoRA

LoRA augments the base model with small, trainable matrices. The original model stays frozen. Key benefits:

  • Less VRAM required.
  • Faster training.
  • Adapters can be swapped easily.
  • Multiple adapters can be combined.

QLoRA goes further by quantizing the base model during training. This makes fine-tuning possible on consumer-grade GPUs.

Preparing training data

Data should be in the right format. A common format for chat models:

{
  "messages": [
    {"role": "system", "content": "You are a support assistant."},
    {"role": "user", "content": "How do I reset my password?"},
    {"role": "assistant", "content": "Go to Settings and select \"Change password\"."}
  ]
}
  • Keep data high-quality and consistent.
  • Aim for at least a few hundred examples, more if possible.
  • Use systematic system prompts for specific response behavior.
  • Respect data privacy; never include sensitive information unprotected.

Training with axolotl or unsloth

Tools like axolotl and unsloth simplify the fine-tuning process.

Example with unsloth

from unsloth import FastLanguageModel
import torch

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/llama-3-8b-bnb-4bit",
    max_seq_length=2048,
    load_in_4bit=True,
)

model = FastLanguageModel.get_peft_model(
    model,
    r=16,
    lora_alpha=16,
    target_modules=["q_proj", "v_proj"],
    lora_dropout=0,
    bias="none",
)

After training, save the adapter.

Converting adapters to GGUF

After training, you usually need to convert the adapter to GGUF format so Ollama can use it:

python convert-lora-to-gguf.py \
  --base-model llama-3-8b \
  --lora adapter \
  --output mein-adapter.gguf

Alternatively, save the merged model directly as GGUF.

Using the model in Ollama

Define the adapter in a Modelfile:

FROM llama3.1
ADAPTER ./mein-adapter.bin

SYSTEM """
You are a specialized assistant for internal support questions.
"""

Create it:

ollama create support-assistent -f Modelfile
ollama run support-assistent

Tips

  • Start with a small dataset and iterate.
  • Adjust learning rate and batch size gradually.
  • Monitor validation loss to catch overfitting early.
  • Test the adapter against questions it hasn’t seen.
  • Explore combining fine-tuning with RAG.

Common pitfalls

  • Too little data: The model won’t train stably.
  • Overfitting: Training too long or on overly specific examples.
  • Wrong format: Ollama expects correct Modelfile syntax.
  • Insufficient VRAM: Use QLoRA or a smaller model.
  • Incompatible adapter: The base model must match.
  • Data leaks: Training data contains sensitive information.

Alternatives to fine-tuning

Before you fine-tune, explore:

  • Prompt engineering: Better-crafted prompts.
  • RAG: Inject external knowledge into context.
  • Model choice: Try a larger model.
  • Modelfile: Adjust system prompts and parameters.

Further reading and resources

FAQ: Fine-tuning Ollama models

Do I need powerful GPUs for fine-tuning? QLoRA enables training on consumer GPUs with 8 to 16 GB of VRAM.

Can I combine multiple LoRA adapters? Yes, either sequentially or by merging them.

Are 100 training examples enough? Usually not. Several hundred to a thousand high-quality examples work better.

Is fine-tuning compliant with data protection regulations? If everything runs locally, your data stays within your network.

How do I test whether it worked? Use a separate validation dataset and unseen test questions.

Sources and further reading

Summary: Fine-tuning Ollama models

Fine-tuning lets you adapt Ollama models to specific tasks and datasets. LoRA and QLoRA make training possible on consumer hardware. What matters most is high-quality training data, the right tools, careful evaluation, and correct integration into Ollama via a Modelfile. Before you fine-tune, check whether prompt engineering or RAG might get you there faster and more cost-effectively.

Back to Blog
Share:

Related Posts