Comparing Ollama Embedding Models
What this article covers
- What embeddings are.
- Which embedding models Ollama offers.
- Differences in dimensions and quality.
- How to use them for RAG and semantic search.
- Recommendations for different use cases.
Introduction: Comparing Ollama embedding models
Embeddings convert text into vectors. These vectors enable semantic search, classification, and RAG. Ollama provides several embedding models that differ in size, quality, and vector dimensions. The right model depends on your application, hardware, and desired accuracy.
This article compares the most common Ollama embedding models and offers recommendations.
Key terms
- Embedding: Vector representation of text.
- Dimension: Length of the vector.
- RAG: Retrieval Augmented Generation.
- Cosine Similarity: Similarity measure between vectors.
- Semantic Search: Search by meaning, not just keywords.
- Token: Unit into which text is divided.
- Quantization: Reduction of model precision.
Available embedding models
| Model | Dimension | Size | Notes |
|---|---|---|---|
| nomic-embed-text | 768 | 137M | Popular, good balance |
| nomic-embed-text-v1.5 | 768 | 137M | Improved version |
| mxbai-embed-large | 1024 | 335M | High quality |
| snowflake-arctic-embed | 768 | 335M | Optimized for retrieval |
| all-minilm | 384 | 22M | Very small, fast |
| bge-m3 | 1024 | 567M | Multilingual |
nomic-embed-text
- Good general-purpose quality.
- 768 dimensions.
- Relatively fast.
- Very common in Ollama examples.
mxbai-embed-large
- 1024 dimensions.
- Often better results for RAG.
- Larger and slower.
- Useful for production RAG systems.
snowflake-arctic-embed
- Optimized for document retrieval.
- Good performance on large datasets.
- 768 or 1024 dimensions depending on variant.
all-minilm
- Very small and fast.
- 384 dimensions.
- For simple tasks and low-resource hardware.
- Less precise than larger models.
bge-m3
- Multilingual strength.
- Supports longer inputs.
- Larger, but flexible.
Choosing a model
| Use case | Recommendation |
|---|---|
| Simple search | all-minilm |
| Standard RAG | nomic-embed-text |
| High quality | mxbai-embed-large |
| Document retrieval | snowflake-arctic-embed |
| Multilingual | bge-m3 |
Pulling models in Ollama
ollama pull nomic-embed-text
ollama pull mxbai-embed-large
Generating embeddings
curl http://localhost:11434/api/embeddings -d '{
"model": "nomic-embed-text",
"prompt": "Docker is a container tool."
}'
Python example
import requests
url = "http://localhost:11434/api/embeddings"
payload = {
"model": "nomic-embed-text",
"prompt": "What is Ollama?"
}
response = requests.post(url, json=payload)
print(len(response.json()["embedding"]))
Token limits matter
Embedding models have token limits:
nomic-embed-text: 2048 tokens.mxbai-embed-large: 512 tokens.bge-m3: 8192 tokens.
Longer texts must be split into chunks.
Comparing quality
You can test quality yourself:
- Generate embeddings for the same texts using different models.
- Check similarity to relevant documents.
- Evaluate relevance of top results.
import numpy as np
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
Memory and speed
| Model | Speed | RAM/VRAM |
|---|---|---|
| all-minilm | very fast | very low |
| nomic-embed-text | fast | low |
| mxbai-embed-large | moderate | moderate |
| bge-m3 | slower | higher |
Tips
- For RAG with Ollama:
nomic-embed-textormxbai-embed-large. - Adjust chunk size to model token limit.
- Use the same model for queries and documents.
- Benchmark on your own data.
- Vector dimension size affects memory usage in vector databases.
Further reading and resources
- BotServ.de Ollama Embeddings
- BotServ.de Local RAG
- BotServ.de Vector Databases
- BotServ.de Ollama Model Recommendations
FAQ: Ollama embedding models
Which is the best embedding model?
For most cases, nomic-embed-text. For highest quality, mxbai-embed-large.
Do I need a specialized model for RAG? Yes, the embedding model and LLM are separate components.
How many dimensions make sense? 384-1024 is sufficient for most applications.
Can I combine embeddings? Not directly. Only models with identical dimensions and similar vector space can be mixed.
What chunk size should I use? Depends on the model’s token limit, typically 256-1024 tokens.
Sources and further reading
- Ollama Embeddings: https://ollama.com/blog/embedding-models
- Nomic Embed: https://docs.nomic.ai/reference/nomic-embed-text-v1
- MTEB Leaderboard: https://huggingface.co/spaces/mteb/leaderboard
Summary: Comparing Ollama embedding models
Ollama offers several embedding models for different requirements. nomic-embed-text is a solid all-rounder, mxbai-embed-large delivers higher quality, and all-minilm is extremely fast and small. Your choice depends on RAG quality needs, available hardware, and speed requirements. What matters most is chunk size, vector dimension, and using the same model consistently for both documents and queries. Testing a few models on your own data will quickly reveal the best fit.


