Fine-tuning Ollama Models
What this article covers
- What fine-tuning is and when it makes sense.
- How LoRA and QLoRA work.
- How to prepare training data.
- How to create an adapter and use it in Ollama.
- Tips, tools, and common pitfalls.
Introduction: Fine-tuning Ollama models
Off-the-shelf models handle many tasks well, but for specialized domains, particular language styles, or proprietary data, their built-in knowledge often falls short. Fine-tuning adapts a base model to a specific purpose. With Ollama, you can run customized models locally by augmenting them with LoRA adapters. This keeps your data private and lets you steer the model precisely toward your needs.
This article walks through the fundamentals of fine-tuning, data preparation, and integration with Ollama.
Key concepts
- Fine-tuning: Further training an already-trained model.
- LoRA: Low-Rank Adaptation, an efficient way to adapt a model with minimal memory overhead.
- QLoRA: Quantized LoRA, even more memory-efficient.
- Adapter: Additional weights that augment a base model.
- Training data: Input-output pairs for learning.
- Evaluation data: Data used to test the trained model.
- Overfitting: The model memorizes training data instead of generalizing.
When fine-tuning makes sense
- You need domain-specific terminology or language patterns.
- Your own knowledge should inform the model’s responses.
- You want a particular response format or tone.
- Simple prompt engineering no longer suffices.
- RAG helps, but the model needs deeper, more specialized capabilities.
If you have limited data or compute, RAG is often the better first step.
LoRA and QLoRA
LoRA augments the base model with small, trainable matrices. The original model stays frozen. Key benefits:
- Less VRAM required.
- Faster training.
- Adapters can be swapped easily.
- Multiple adapters can be combined.
QLoRA goes further by quantizing the base model during training. This makes fine-tuning possible on consumer-grade GPUs.
Preparing training data
Data should be in the right format. A common format for chat models:
{
"messages": [
{"role": "system", "content": "You are a support assistant."},
{"role": "user", "content": "How do I reset my password?"},
{"role": "assistant", "content": "Go to Settings and select \"Change password\"."}
]
}
- Keep data high-quality and consistent.
- Aim for at least a few hundred examples, more if possible.
- Use systematic system prompts for specific response behavior.
- Respect data privacy; never include sensitive information unprotected.
Training with axolotl or unsloth
Tools like axolotl and unsloth simplify the fine-tuning process.
Example with unsloth
from unsloth import FastLanguageModel
import torch
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/llama-3-8b-bnb-4bit",
max_seq_length=2048,
load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(
model,
r=16,
lora_alpha=16,
target_modules=["q_proj", "v_proj"],
lora_dropout=0,
bias="none",
)
After training, save the adapter.
Converting adapters to GGUF
After training, you usually need to convert the adapter to GGUF format so Ollama can use it:
python convert-lora-to-gguf.py \
--base-model llama-3-8b \
--lora adapter \
--output mein-adapter.gguf
Alternatively, save the merged model directly as GGUF.
Using the model in Ollama
Define the adapter in a Modelfile:
FROM llama3.1
ADAPTER ./mein-adapter.bin
SYSTEM """
You are a specialized assistant for internal support questions.
"""
Create it:
ollama create support-assistent -f Modelfile
ollama run support-assistent
Tips
- Start with a small dataset and iterate.
- Adjust learning rate and batch size gradually.
- Monitor validation loss to catch overfitting early.
- Test the adapter against questions it hasn’t seen.
- Explore combining fine-tuning with RAG.
Common pitfalls
- Too little data: The model won’t train stably.
- Overfitting: Training too long or on overly specific examples.
- Wrong format: Ollama expects correct Modelfile syntax.
- Insufficient VRAM: Use QLoRA or a smaller model.
- Incompatible adapter: The base model must match.
- Data leaks: Training data contains sensitive information.
Alternatives to fine-tuning
Before you fine-tune, explore:
- Prompt engineering: Better-crafted prompts.
- RAG: Inject external knowledge into context.
- Model choice: Try a larger model.
- Modelfile: Adjust system prompts and parameters.
Further reading and resources
- BotServ.de Ollama Modelfiles
- BotServ.de Ollama Model lifecycle
- BotServ.de RAG knowledge base
- BotServ.de Ollama performance
FAQ: Fine-tuning Ollama models
Do I need powerful GPUs for fine-tuning? QLoRA enables training on consumer GPUs with 8 to 16 GB of VRAM.
Can I combine multiple LoRA adapters? Yes, either sequentially or by merging them.
Are 100 training examples enough? Usually not. Several hundred to a thousand high-quality examples work better.
Is fine-tuning compliant with data protection regulations? If everything runs locally, your data stays within your network.
How do I test whether it worked? Use a separate validation dataset and unseen test questions.
Sources and further reading
- Unsloth: https://github.com/unslothai/unsloth
- Axolotl: https://github.com/OpenAccess-AI-Collective/axolotl
- LoRA Paper: https://arxiv.org/abs/2106.09685
- QLoRA Paper: https://arxiv.org/abs/2305.14314
Summary: Fine-tuning Ollama models
Fine-tuning lets you adapt Ollama models to specific tasks and datasets. LoRA and QLoRA make training possible on consumer hardware. What matters most is high-quality training data, the right tools, careful evaluation, and correct integration into Ollama via a Modelfile. Before you fine-tune, check whether prompt engineering or RAG might get you there faster and more cost-effectively.


