Skip to content
BotServBotServ
OllamaGGUFImportModelfileQuantization

Import GGUF Models into Ollama

Use custom GGUF models in Ollama. Modelfiles, quantization, parameters, and import steps.

S

schutzgeist

3 min read
Import GGUF Models into Ollama

Importing GGUF Models into Ollama

What this article covers

  • What the GGUF format is.
  • How to import a GGUF model into Ollama.
  • Building an Ollama Modelfile.
  • Quantization and parameters.
  • Common errors and solutions.

Introduction: Importing GGUF models into Ollama

Ollama works with GGUF files. If you want to use your own models or ones downloaded from Hugging Face, you can convert them to GGUF format and import them into Ollama. The import happens through a Modelfile, which points to the GGUF file and sets parameters like temperature, context length, and system prompt. This lets you run specialized models, fine-tuned variants, or the latest quantizations locally.

This article walks through the complete GGUF-to-Ollama import process.

Key terminology

  • GGUF: GPT-Generated Unified Format, a binary format for LLMs.
  • Modelfile: Ollama’s recipe for a model.
  • FROM: Reference to a GGUF file or existing model.
  • PARAMETER: Runtime parameters like temperature.
  • TEMPLATE: Prompt formatting.
  • SYSTEM: System prompt.
  • Quantization: Reducing model precision.
  • Adapter: LoRA adapters for fine-tuning.

Requirements

  • Ollama installed.
  • GGUF file available.
  • Sufficient storage and RAM/VRAM.

Sources for GGUF models

  • Hugging Face TheBloke repository.
  • Individual model pages with GGUF files.
  • Custom conversion using convert.py or llama.cpp.

A simple Modelfile

FROM ./mein-modell.Q4_K_M.gguf

PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER top_k 40
PARAMETER num_ctx 4096

SYSTEM "Hilfreicher Assistent"

Save it as Modelfile in the same directory as the GGUF file.

Creating the model

ollama create mein-modell -f Modelfile

Testing

ollama run mein-modell

Using an existing Ollama model as a base

Instead of a GGUF file, you can use an Ollama model as your base:

FROM llama3.1

PARAMETER temperature 0.5
SYSTEM "Fachspezifischer Assistent"

Customizing the template

The template controls how prompt, system message, and response are assembled:

TEMPLATE """{{ if .System }}<|start_header_id|>system<|end_header_id|>
{{ .System }}<|eot_id|>{{ end }}{{ if .Prompt }}<|start_header_id|>user<|end_header_id|>
{{ .Prompt }}<|eot_id|>{{ end }}<|start_header_id|>assistant<|end_header_id|>
{{ .Response }}<|eot_id|>"""

This example follows the Llama 3 format.

Using LoRA adapters

FROM llama3.1
ADAPTER ./mein-lora.bin
PARAMETER temperature 0.7

Choosing quantization

QuantSizeQuality
Q4_0smallacceptable
Q4_K_Msmallgood
Q5_K_Mmediumvery good
Q6_Klargerhigh
Q8_0largevery high

Lower quantization reduces VRAM usage but degrades quality.

Parameters explained

  • temperature: Balances creativity versus determinism.
  • top_p: Nucleus sampling.
  • top_k: Limits the top-K tokens.
  • num_ctx: Context length.
  • num_predict: Maximum response length.
  • seed: Reproducibility.

Troubleshooting

  • File not found: Check the path in your Modelfile.
  • Invalid GGUF: Verify the version and re-convert if needed.
  • Insufficient RAM: Choose a smaller quantization.
  • Wrong template: Adjust formatting to match your model.
  • Parameters ignored: Recreate the model.

Tips

  • Version control your Modelfiles.
  • Use relative paths to GGUF files.
  • Check available disk space before importing.
  • Inspect existing models with ollama show --modelfile.
  • Use a consistent naming scheme.

Further reading and resources

FAQ: GGUF imports

Where do I find GGUF files? Hugging Face is the most common source.

Do I need to restart Ollama? No, running ollama create is sufficient.

Can I have multiple quantizations? Yes, create a separate model for each quantization.

What if the model won’t start? Check RAM/VRAM, reduce quantization, and review the logs.

Can I use ChatML templates? Yes, customize the template in your Modelfile.

References and further reading

Summary: Importing GGUF models into Ollama

Importing your own GGUF models into Ollama opens access to specialized and up-to-date models. A Modelfile defines the base, parameters, template, and system prompt. Quantization, storage space, and appropriate templates are the key success factors. By versioning your Modelfiles and carefully checking for errors, you can integrate any GGUF file into your local Ollama environment and significantly expand your available model options.

Back to Blog
Share:

Related Posts