Quantization in Practice with Ollama
What this article covers
- Available quantization methods.
- Converting models to GGUF using llama.cpp.
- Differences between Q4_K_M, Q5_K_M, Q8_0, and other approaches.
- Importing custom GGUF files into Ollama.
- Balancing quality, speed, and memory usage.
Introduction: Quantization in Practice with Ollama
Ollama ships with many models already quantized at different levels. Sometimes you want to bring a specific model from Hugging Face into the right format yourself, perhaps because the standard tags don’t fit or you need precise control over file size. You can convert the model to GGUF and then quantize it. The method you choose determines both your memory savings and quality loss.
This guide walks through the key quantization levels and the practical workflow with llama.cpp.
Key Terminology
- Quantization: Reducing the bit depth of model weights.
- GGUF: File format used by llama.cpp for storing models.
- Q4_K_M: 4-bit quantization, K-quant medium, a solid compromise.
- Q5_K_M: 5-bit quantization, higher quality than Q4.
- Q8_0: 8-bit quantization, nearly original quality.
- F16: Half precision, 16 bits per parameter.
- FP32: Full 32-bit floating point precision.
- imatrix: Importance matrix for improved quantization.
- KLD: Perplexity-based quality assessment method.
Why Quantize?
- Larger models fit into less VRAM.
- Run multiple models simultaneously without swapping.
- Faster load times.
- Enable CPUs to run models at all.
- Reduce disk storage footprint.
The trade-off is potential quality loss. With 4-bit quantization, many models show minimal degradation, though complex tasks may suffer noticeably.
Common Quantization Levels
| Level | Bits | Memory | Quality | Use Case |
|---|---|---|---|---|
| Q2_K | 2 | very low | poor | Emergency CPU mode |
| Q3_K_S / Q3_K_M | 3 | low | fair | Severely limited VRAM |
| Q4_K_M | 4 | good | good | Standard for consumer |
| Q5_K_M | 5 | higher | very good | Better fidelity |
| Q8_0 | 8 | high | nearly identical | When space allows |
| F16 | 16 | very high | original | Fine-tuning or training |
Prerequisites
Start by building llama.cpp or using unsloth:
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make
Precompiled binaries and Docker images are also available.
Converting Models to GGUF
For Hugging Face models, use convert_hf_to_gguf.py:
python convert_hf_to_gguf.py /path/to/model \
--outfile my-model-f16.gguf \
--outtype f16
This produces an F16 GGUF file that hasn’t yet been quantized.
Quantizing with llama-quantize
./llama-quantize my-model-f16.gguf my-model-q4_k_m.gguf Q4_K_M
Create additional variants:
./llama-quantize my-model-f16.gguf my-model-q5_k_m.gguf Q5_K_M
./llama-quantize my-model-f16.gguf my-model-q8_0.gguf Q8_0
Using an Importance Matrix
For better Q4 and Q5 results, compute an imatrix first:
./llama-imatrix -m my-model-f16.gguf -f training-data.txt --output-file my-model.imatrix
./llama-quantize --imatrix my-model.imatrix my-model-f16.gguf my-model-q4_k_m.gguf Q4_K_M
An imatrix helps preserve the precision of important weights during quantization.
Importing a Model into Ollama
Create a Modelfile:
FROM ./my-model-q4_k_m.gguf
SYSTEM """
You are a precise assistant.
"""
PARAMETER temperature 0.7
Then import it:
ollama create my-model -f Modelfile
Offering Multiple Quantizations
You can import the same model at different levels:
ollama create my-model:q4 -f Modelfile.q4
ollama create my-model:q5 -f Modelfile.q5
ollama create my-model:q8 -f Modelfile.q8
Assessing Quality
After quantizing, test your model on known tasks. Key areas include:
- Summarization
- Coding tasks
- Math
- Multilingual reasoning
- Logical inference
Compare against the F16 original to see how much quality degrades.
Size and Speed
- Smaller quantized models are faster because they need less memory bandwidth.
- Q2_K and Q3_K_S can cause severe quality loss.
- Q4_K_M is the standard consumer-GPU compromise.
- Q8_0 is nearly lossless but requires significantly more memory.
Tips
- Use Q4_K_M for everyday work.
- Use Q5_K_M for coding or complex reasoning.
- Use Q8_0 if you have enough VRAM.
- Reserve F16 for training or special cases.
- Compute an imatrix for important models.
- Always compare results to the original.
Common Pitfalls
- Wrong format: Hugging Face models must be converted to GGUF first.
- Unsupported architecture: Not every model is compatible with llama.cpp.
- Over-quantization: Q2_K or Q3_K_S often destroys usability.
- Missing tokenizer files: Conversion requires complete model artifacts.
- Still out of VRAM: Model doesn’t fit even after quantization.
- Corrupt GGUF: Verify checksums or perform a test run.
Further Reading and Resources
- BotServ.de Ollama Modelfiles
- BotServ.de Ollama Performance
- BotServ.de Quantization Basics
- BotServ.de Quantization Calculator
FAQ: Quantization in Practice
Which quantization is the best compromise? Q4_K_M for most use cases.
Do you lose much quality? Usually minimal with Q4_K_M, almost nothing with Q8_0.
Can I use Hugging Face models with Ollama? Yes, convert them to GGUF and import them into Ollama.
Do I need Linux? llama.cpp runs on Windows and macOS too, but Linux is most common.
How long does quantization take? Minutes to hours depending on model size and CPU.
Sources and Further Reading
- llama.cpp: https://github.com/ggerganov/llama.cpp
- GGUF Format: https://github.com/ggerganov/ggml/blob/master/docs/gguf.md
- TheBloke on Hugging Face: https://huggingface.co/TheBloke
Summary: Quantization in Practice with Ollama
Quantization is essential for running large models on consumer hardware. With llama.cpp, you can convert Hugging Face models to GGUF and quantize them at levels like Q4_K_M, Q5_K_M, or Q8_0. For the best results with minimal memory footprint, compute an importance matrix. After importing via an Ollama Modelfile, your quantized model works like any other Ollama model. The key is balancing quality against memory constraints and choosing the right level for your hardware and workload.


