Importing GGUF Models into Ollama
What this article covers
- What the GGUF format is.
- How to import a GGUF model into Ollama.
- Building an Ollama Modelfile.
- Quantization and parameters.
- Common errors and solutions.
Introduction: Importing GGUF models into Ollama
Ollama works with GGUF files. If you want to use your own models or ones downloaded from Hugging Face, you can convert them to GGUF format and import them into Ollama. The import happens through a Modelfile, which points to the GGUF file and sets parameters like temperature, context length, and system prompt. This lets you run specialized models, fine-tuned variants, or the latest quantizations locally.
This article walks through the complete GGUF-to-Ollama import process.
Key terminology
- GGUF: GPT-Generated Unified Format, a binary format for LLMs.
- Modelfile: Ollama’s recipe for a model.
- FROM: Reference to a GGUF file or existing model.
- PARAMETER: Runtime parameters like temperature.
- TEMPLATE: Prompt formatting.
- SYSTEM: System prompt.
- Quantization: Reducing model precision.
- Adapter: LoRA adapters for fine-tuning.
Requirements
- Ollama installed.
- GGUF file available.
- Sufficient storage and RAM/VRAM.
Sources for GGUF models
- Hugging Face TheBloke repository.
- Individual model pages with GGUF files.
- Custom conversion using
convert.pyorllama.cpp.
A simple Modelfile
FROM ./mein-modell.Q4_K_M.gguf
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER top_k 40
PARAMETER num_ctx 4096
SYSTEM "Hilfreicher Assistent"
Save it as Modelfile in the same directory as the GGUF file.
Creating the model
ollama create mein-modell -f Modelfile
Testing
ollama run mein-modell
Using an existing Ollama model as a base
Instead of a GGUF file, you can use an Ollama model as your base:
FROM llama3.1
PARAMETER temperature 0.5
SYSTEM "Fachspezifischer Assistent"
Customizing the template
The template controls how prompt, system message, and response are assembled:
TEMPLATE """{{ if .System }}<|start_header_id|>system<|end_header_id|>
{{ .System }}<|eot_id|>{{ end }}{{ if .Prompt }}<|start_header_id|>user<|end_header_id|>
{{ .Prompt }}<|eot_id|>{{ end }}<|start_header_id|>assistant<|end_header_id|>
{{ .Response }}<|eot_id|>"""
This example follows the Llama 3 format.
Using LoRA adapters
FROM llama3.1
ADAPTER ./mein-lora.bin
PARAMETER temperature 0.7
Choosing quantization
| Quant | Size | Quality |
|---|---|---|
| Q4_0 | small | acceptable |
| Q4_K_M | small | good |
| Q5_K_M | medium | very good |
| Q6_K | larger | high |
| Q8_0 | large | very high |
Lower quantization reduces VRAM usage but degrades quality.
Parameters explained
temperature: Balances creativity versus determinism.top_p: Nucleus sampling.top_k: Limits the top-K tokens.num_ctx: Context length.num_predict: Maximum response length.seed: Reproducibility.
Troubleshooting
- File not found: Check the path in your Modelfile.
- Invalid GGUF: Verify the version and re-convert if needed.
- Insufficient RAM: Choose a smaller quantization.
- Wrong template: Adjust formatting to match your model.
- Parameters ignored: Recreate the model.
Tips
- Version control your Modelfiles.
- Use relative paths to GGUF files.
- Check available disk space before importing.
- Inspect existing models with
ollama show --modelfile. - Use a consistent naming scheme.
Further reading and resources
- BotServ.de Ollama Modelfiles
- BotServ.de Ollama Quantization in Practice
- BotServ.de Ollama Fine-tuning
- BotServ.de Ollama Commands
FAQ: GGUF imports
Where do I find GGUF files? Hugging Face is the most common source.
Do I need to restart Ollama?
No, running ollama create is sufficient.
Can I have multiple quantizations? Yes, create a separate model for each quantization.
What if the model won’t start? Check RAM/VRAM, reduce quantization, and review the logs.
Can I use ChatML templates? Yes, customize the template in your Modelfile.
References and further reading
- Ollama Modelfile: https://github.com/ollama/ollama/blob/main/docs/modelfile.md
- GGUF: https://github.com/ggerganov/ggml/blob/master/docs/gguf.md
- Hugging Face: https://huggingface.co/
Summary: Importing GGUF models into Ollama
Importing your own GGUF models into Ollama opens access to specialized and up-to-date models. A Modelfile defines the base, parameters, template, and system prompt. Quantization, storage space, and appropriate templates are the key success factors. By versioning your Modelfiles and carefully checking for errors, you can integrate any GGUF file into your local Ollama environment and significantly expand your available model options.


