Embedding Models Compared
What This Article Covers
- What embedding models are and how they work.
- nomic-embed-text, multilingual-e5, bge-m3, and others side by side.
- Which model for which language and use case.
- Hardware requirements and speed.
- Practical examples for RAG and semantic search.
Introduction: Understanding Embedding Models
Embedding models convert text into numbers (vectors). Similar texts produce similar vectors. This enables semantic search: instead of matching keywords, the system finds texts with similar meaning. Embeddings form the foundation for RAG, semantic search, and document clustering.
This article is for anyone choosing embedding models for RAG or search. For foundational concepts, see Embeddings and Local RAG.
Why Do You Need Embedding Models?
Imagine searching for “vacation request” and wanting to find documents containing “request time off” or “leave application,” even though the words differ. Embeddings understand meaning, not just words. That’s the basis for RAG and intelligent search.
Embedding Models Explained Briefly
Embedding models convert text into vectors (for example, 768 numbers). Similar texts have similar vectors. nomic-embed-text is a solid general-purpose choice, multilingual-e5 handles multiple languages, and bge-m3 works well with lengthy documents.
The core idea is simple: text → numbers → find similarity.
Who This Article Is For
- RAG developers needing embeddings for document search.
- Search engine builders implementing semantic search.
- Data analysts clustering and classifying texts.
- Self-hosters creating embeddings locally.
Some RAG background is helpful.
Key Terms
- Embedding - Text converted to a vector. Useful for: semantic search.
- RAG - Retrieval-Augmented Generation. Useful for: the primary use case.
- Vector database - Storage for embeddings. Useful for: search.
- Semantic search - Meaning over keywords. Useful for: better results.
- Ollama - Model server. Useful for: embeddings.
- Cosine similarity - Similarity measure. Useful for: comparisons.
How Embeddings Work
Text: "Der Hund spielt im Garten"
│
▼
Embedding-Modell
│
▼
Vektor: [0.23, -0.45, 0.78, ..., 0.12] (768 Dimensionen)
Ähnlicher Text: "Ein Hund spielt draußen"
│
▼
Vektor: [0.21, -0.43, 0.81, ..., 0.15] (ähnlich!)
Unähnlicher Text: "Die Katze schläft auf dem Sofa"
│
▼
Vektor: [-0.67, 0.89, -0.23, ..., 0.45] (unähnlich)
Models at a Glance
| Model | Developer | Dimensions | Languages | Max Length | Best For |
|---|---|---|---|---|---|
| nomic-embed-text | Nomic AI | 768 | English | 8K | General RAG |
| multilingual-e5 | Microsoft | 1024 | 100+ | 512 | Multilingual RAG |
| bge-m3 | BAAI | 1024 | 100+ | 8K | Long documents |
| bge-large-en | BAAI | 1024 | English | 512 | English texts |
| gte-Qwen2 | Alibaba | 1536 | Multilingual | 32K | Very long documents |
| mxbai-embed-large | Mixedbread | 1024 | English | 512 | English RAG |
Detailed Comparison
1. nomic-embed-text (Nomic AI)
Strengths:
- Default embedding for Ollama
- Strong quality for English text
- Fast and efficient
- Well documented
Weaknesses:
- English only (suboptimal for German)
- 768 dimensions (fewer than others)
Recommendation: Use nomic-embed-text for English documents.
2. multilingual-e5 (Microsoft)
Strengths:
- 100+ languages including German
- Excellent quality
- Standard for multilingual applications
- multilingual-e5-large is the best version
Weaknesses:
- Only 512 token context (short texts)
- Larger model size
Recommendation: Use multilingual-e5 for German and multilingual documents.
3. bge-m3 (BAAI)
Strengths:
- Multilingual (100+ languages)
- 8K context (long documents)
- Excellent quality
- Hybrid: dense + sparse + multi-vector
Weaknesses:
- Larger and slower
- More complex (three modes)
Recommendation: Use bge-m3 for long, multilingual documents.
4. gte-Qwen2 (Alibaba)
Strengths:
- 32K context (very long documents)
- Multilingual
- Excellent quality
- Qwen-based
Weaknesses:
- 1536 dimensions (more memory)
- Less widely adopted
Recommendation: Use gte-qwen2 for very long documents.
Language Support Comparison
| Model | German | English | Multilingual | Note |
|---|---|---|---|---|
| nomic-embed-text | ⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐ | English only |
| multilingual-e5 | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | 100+ languages |
| bge-m3 | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | 100+ languages |
| gte-Qwen2 | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | Multilingual |
Context Length Comparison
| Model | Max Tokens | For What |
|---|---|---|
| multilingual-e5 | 512 | Short texts, sentences |
| bge-large-en | 512 | Short texts |
| nomic-embed-text | 8K | Medium documents |
| bge-m3 | 8K | Long documents |
| gte-Qwen2 | 32K | Very long documents |
Practical Example: Embeddings with Ollama
import requests
# Create an embedding
def get_embedding(text, model="nomic-embed-text"):
response = requests.post(
"http://ollama:11434/api/embeddings",
json={
"model": model,
"prompt": text
}
)
return response.json()["embedding"]
# Example
text = "Der Hund spielt im Garten"
embedding = get_embedding(text)
print(f"Dimensionen: {len(embedding)}")
# Output: Dimensionen: 768
Practical Example: RAG Pipeline
import chromadb
import requests
# 1. Embed and store documents
chroma = chromadb.HttpClient(host="chromadb", port=8000)
collection = chroma.get_or_create_collection("dokumente")
documents = [
"Der Hund spielt im Garten.",
"Katzen schlafen gerne auf dem Sofa.",
"Vögel fliegen am Himmel."
]
for i, doc in enumerate(documents):
embedding = get_embedding(doc)
collection.add(
ids=[f"doc_{i}"],
documents=[doc],
embeddings=[embedding]
)
# 2. Semantic search
query = "Welches Tier spielt draußen?"
query_embedding = get_embedding(query)
results = collection.query(
query_embeddings=[query_embedding],
n_results=1
)
print(results["documents"][0])
# Output: ["Der Hund spielt im Garten."]
Practical Example: German Documents
# multilingual-e5 for German documents
def get_german_embedding(text):
response = requests.post(
"http://ollama:11434/api/embeddings",
json={
"model": "multilingual-e5",
"prompt": text
}
)
return response.json()["embedding"]
# German documents
docs = [
"Der Mitarbeiter beantragt Urlaub.",
"Die Rechnung ist fällig.",
"Das Meeting findet morgen statt."
]
for doc in docs:
emb = get_german_embedding(doc)
# Store in vector database
Security Considerations
- Embeddings are not reversible: you cannot reconstruct the original text from a vector. However, they can still contain sensitive information.
- Local embeddings: With Ollama, all embeddings remain local. With cloud APIs, texts are sent externally.
- Bias: Embedding models have biases. Test thoroughly for critical applications.
- Dimensions: Higher dimensions mean more memory, but not always better results.
Common Pitfalls
- Wrong model for the language: nomic-embed-text for German documents = suboptimal. Use multilingual-e5 instead.
- Context length too short: Long documents need chunking. Use bge-m3 or gte-Qwen2 for extended texts.
- Embedding model for generation: Embedding models don’t generate text. They only create vectors.
- Similarity comparison without normalization: Cosine similarity requires normalized vectors.
- Too many dimensions: 1536 dimensions consume more memory and aren’t always better.
Further Reading
- Embeddings - Fundamentals.
- Embedding Models - For RAG.
- Local RAG - RAG pipeline.
- Vector Databases - Storage.
- Ollama - Model server.
- Chunking - Text splitting.
- Hybrid Search - Combined search.
Key Takeaways:
- Embeddings convert text into vectors for semantic search.
- nomic-embed-text = standard for English.
- multilingual-e5 = best for German and multilingual support.
- bge-m3 = best for long documents.
- gte-Qwen2 = for very long documents (32K).
- For German RAG: multilingual-e5 or bge-m3.
FAQ
What is an embedding model?
Which embedding model for German?
What do dimensions mean?
How long can texts be?
Embedding model or LLM?
Can I create embeddings with Ollama?
What is hybrid search?
What do embeddings cost?
Sources and Further Reading
- nomic-embed-text - Ollama model.
- multilingual-e5 - Microsoft model.
- bge-m3 - BAAI model.
- Ollama Embeddings - Embeddings API.
- MTEB Leaderboard - Embedding benchmarks.


