Skip to content
BotServBotServ
Embedding Modelsnomic-embed-textmultilingual-e5bge-m3RAGSemantic Search

Embedding Models Comparison

Compare embedding models for local AI: nomic-embed-text, multilingual-e5, bge-m3 for RAG and semantic search.

S

schutzgeist

5 min read
Embedding Models Comparison

Embedding Models Compared

What This Article Covers

  • What embedding models are and how they work.
  • nomic-embed-text, multilingual-e5, bge-m3, and others side by side.
  • Which model for which language and use case.
  • Hardware requirements and speed.
  • Practical examples for RAG and semantic search.

Introduction: Understanding Embedding Models

Embedding models convert text into numbers (vectors). Similar texts produce similar vectors. This enables semantic search: instead of matching keywords, the system finds texts with similar meaning. Embeddings form the foundation for RAG, semantic search, and document clustering.

This article is for anyone choosing embedding models for RAG or search. For foundational concepts, see Embeddings and Local RAG.

Why Do You Need Embedding Models?

Imagine searching for “vacation request” and wanting to find documents containing “request time off” or “leave application,” even though the words differ. Embeddings understand meaning, not just words. That’s the basis for RAG and intelligent search.

Embedding Models Explained Briefly

Embedding models convert text into vectors (for example, 768 numbers). Similar texts have similar vectors. nomic-embed-text is a solid general-purpose choice, multilingual-e5 handles multiple languages, and bge-m3 works well with lengthy documents.

The core idea is simple: text → numbers → find similarity.

Who This Article Is For

  • RAG developers needing embeddings for document search.
  • Search engine builders implementing semantic search.
  • Data analysts clustering and classifying texts.
  • Self-hosters creating embeddings locally.

Some RAG background is helpful.

Key Terms

  • Embedding - Text converted to a vector. Useful for: semantic search.
  • RAG - Retrieval-Augmented Generation. Useful for: the primary use case.
  • Vector database - Storage for embeddings. Useful for: search.
  • Semantic search - Meaning over keywords. Useful for: better results.
  • Ollama - Model server. Useful for: embeddings.
  • Cosine similarity - Similarity measure. Useful for: comparisons.

How Embeddings Work

Text: "Der Hund spielt im Garten"
         │
         ▼
Embedding-Modell
         │
         ▼
Vektor: [0.23, -0.45, 0.78, ..., 0.12]  (768 Dimensionen)

Ähnlicher Text: "Ein Hund spielt draußen"
         │
         ▼
Vektor: [0.21, -0.43, 0.81, ..., 0.15]  (ähnlich!)

Unähnlicher Text: "Die Katze schläft auf dem Sofa"
         │
         ▼
Vektor: [-0.67, 0.89, -0.23, ..., 0.45]  (unähnlich)

Models at a Glance

ModelDeveloperDimensionsLanguagesMax LengthBest For
nomic-embed-textNomic AI768English8KGeneral RAG
multilingual-e5Microsoft1024100+512Multilingual RAG
bge-m3BAAI1024100+8KLong documents
bge-large-enBAAI1024English512English texts
gte-Qwen2Alibaba1536Multilingual32KVery long documents
mxbai-embed-largeMixedbread1024English512English RAG

Detailed Comparison

1. nomic-embed-text (Nomic AI)

Strengths:

  • Default embedding for Ollama
  • Strong quality for English text
  • Fast and efficient
  • Well documented

Weaknesses:

  • English only (suboptimal for German)
  • 768 dimensions (fewer than others)

Recommendation: Use nomic-embed-text for English documents.

2. multilingual-e5 (Microsoft)

Strengths:

  • 100+ languages including German
  • Excellent quality
  • Standard for multilingual applications
  • multilingual-e5-large is the best version

Weaknesses:

  • Only 512 token context (short texts)
  • Larger model size

Recommendation: Use multilingual-e5 for German and multilingual documents.

3. bge-m3 (BAAI)

Strengths:

  • Multilingual (100+ languages)
  • 8K context (long documents)
  • Excellent quality
  • Hybrid: dense + sparse + multi-vector

Weaknesses:

  • Larger and slower
  • More complex (three modes)

Recommendation: Use bge-m3 for long, multilingual documents.

4. gte-Qwen2 (Alibaba)

Strengths:

  • 32K context (very long documents)
  • Multilingual
  • Excellent quality
  • Qwen-based

Weaknesses:

  • 1536 dimensions (more memory)
  • Less widely adopted

Recommendation: Use gte-qwen2 for very long documents.

Language Support Comparison

ModelGermanEnglishMultilingualNote
nomic-embed-text⭐⭐⭐⭐⭐⭐⭐⭐English only
multilingual-e5⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐100+ languages
bge-m3⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐100+ languages
gte-Qwen2⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐Multilingual

Context Length Comparison

ModelMax TokensFor What
multilingual-e5512Short texts, sentences
bge-large-en512Short texts
nomic-embed-text8KMedium documents
bge-m38KLong documents
gte-Qwen232KVery long documents

Practical Example: Embeddings with Ollama

import requests

# Create an embedding
def get_embedding(text, model="nomic-embed-text"):
    response = requests.post(
        "http://ollama:11434/api/embeddings",
        json={
            "model": model,
            "prompt": text
        }
    )
    return response.json()["embedding"]

# Example
text = "Der Hund spielt im Garten"
embedding = get_embedding(text)
print(f"Dimensionen: {len(embedding)}")
# Output: Dimensionen: 768

Practical Example: RAG Pipeline

import chromadb
import requests

# 1. Embed and store documents
chroma = chromadb.HttpClient(host="chromadb", port=8000)
collection = chroma.get_or_create_collection("dokumente")

documents = [
    "Der Hund spielt im Garten.",
    "Katzen schlafen gerne auf dem Sofa.",
    "Vögel fliegen am Himmel."
]

for i, doc in enumerate(documents):
    embedding = get_embedding(doc)
    collection.add(
        ids=[f"doc_{i}"],
        documents=[doc],
        embeddings=[embedding]
    )

# 2. Semantic search
query = "Welches Tier spielt draußen?"
query_embedding = get_embedding(query)

results = collection.query(
    query_embeddings=[query_embedding],
    n_results=1
)

print(results["documents"][0])
# Output: ["Der Hund spielt im Garten."]

Practical Example: German Documents

# multilingual-e5 for German documents
def get_german_embedding(text):
    response = requests.post(
        "http://ollama:11434/api/embeddings",
        json={
            "model": "multilingual-e5",
            "prompt": text
        }
    )
    return response.json()["embedding"]

# German documents
docs = [
    "Der Mitarbeiter beantragt Urlaub.",
    "Die Rechnung ist fällig.",
    "Das Meeting findet morgen statt."
]

for doc in docs:
    emb = get_german_embedding(doc)
    # Store in vector database

Security Considerations

  • Embeddings are not reversible: you cannot reconstruct the original text from a vector. However, they can still contain sensitive information.
  • Local embeddings: With Ollama, all embeddings remain local. With cloud APIs, texts are sent externally.
  • Bias: Embedding models have biases. Test thoroughly for critical applications.
  • Dimensions: Higher dimensions mean more memory, but not always better results.

Common Pitfalls

  • Wrong model for the language: nomic-embed-text for German documents = suboptimal. Use multilingual-e5 instead.
  • Context length too short: Long documents need chunking. Use bge-m3 or gte-Qwen2 for extended texts.
  • Embedding model for generation: Embedding models don’t generate text. They only create vectors.
  • Similarity comparison without normalization: Cosine similarity requires normalized vectors.
  • Too many dimensions: 1536 dimensions consume more memory and aren’t always better.

Further Reading

Key Takeaways:

  • Embeddings convert text into vectors for semantic search.
  • nomic-embed-text = standard for English.
  • multilingual-e5 = best for German and multilingual support.
  • bge-m3 = best for long documents.
  • gte-Qwen2 = for very long documents (32K).
  • For German RAG: multilingual-e5 or bge-m3.

FAQ

What is an embedding model?

A model that transforms text into numbers (vectors). Similar texts produce similar vectors. This enables semantic search: meaning instead of keywords.

Which embedding model for German?

multilingual-e5 or bge-m3. Both support German very well. nomic-embed-text is optimized for English only.

What do dimensions mean?

The number of values in the vector. 768 (nomic), 1024 (e5, bge), 1536 (gte). More dimensions consume more memory, but don’t always improve quality.

How long can texts be?

multilingual-e5: 512 tokens. bge-m3: 8K tokens. gte-Qwen2: 32K tokens. Longer texts must be chunked.

Embedding model or LLM?

Different purposes. Embedding models create vectors for search. LLMs generate text. For RAG you need both: embeddings for retrieval, LLM for generation.

Can I create embeddings with Ollama?

Yes, via the Ollama Embeddings API: POST /api/embeddings with model and text. Supports nomic-embed-text, multilingual-e5, and more.

What is hybrid search?

A combination of semantic search (embeddings) and keyword search (BM25). Results are merged for better matches. bge-m3 supports both.

What do embeddings cost?

Locally: hardware costs only. Ollama + multilingual-e5 are free. Cloud APIs (OpenAI, Cohere) charge per million tokens.

Sources and Further Reading

Back to Blog
Share:

Related Posts