Skip to content
BotServBotServ
OllamaEmbeddingsnomicmxbaisnowflake

Compare Ollama Embedding Models

Compare embedding models for Ollama and RAG: nomic, mxbai, snowflake, MiniLM. Dimensions, performance, use cases.

S

schutzgeist

2 min read
Compare Ollama Embedding Models

Comparing Ollama Embedding Models

What this article covers

  • What embeddings are.
  • Which embedding models Ollama offers.
  • Differences in dimensions and quality.
  • How to use them for RAG and semantic search.
  • Recommendations for different use cases.

Introduction: Comparing Ollama embedding models

Embeddings convert text into vectors. These vectors enable semantic search, classification, and RAG. Ollama provides several embedding models that differ in size, quality, and vector dimensions. The right model depends on your application, hardware, and desired accuracy.

This article compares the most common Ollama embedding models and offers recommendations.

Key terms

  • Embedding: Vector representation of text.
  • Dimension: Length of the vector.
  • RAG: Retrieval Augmented Generation.
  • Cosine Similarity: Similarity measure between vectors.
  • Semantic Search: Search by meaning, not just keywords.
  • Token: Unit into which text is divided.
  • Quantization: Reduction of model precision.

Available embedding models

ModelDimensionSizeNotes
nomic-embed-text768137MPopular, good balance
nomic-embed-text-v1.5768137MImproved version
mxbai-embed-large1024335MHigh quality
snowflake-arctic-embed768335MOptimized for retrieval
all-minilm38422MVery small, fast
bge-m31024567MMultilingual

nomic-embed-text

  • Good general-purpose quality.
  • 768 dimensions.
  • Relatively fast.
  • Very common in Ollama examples.

mxbai-embed-large

  • 1024 dimensions.
  • Often better results for RAG.
  • Larger and slower.
  • Useful for production RAG systems.

snowflake-arctic-embed

  • Optimized for document retrieval.
  • Good performance on large datasets.
  • 768 or 1024 dimensions depending on variant.

all-minilm

  • Very small and fast.
  • 384 dimensions.
  • For simple tasks and low-resource hardware.
  • Less precise than larger models.

bge-m3

  • Multilingual strength.
  • Supports longer inputs.
  • Larger, but flexible.

Choosing a model

Use caseRecommendation
Simple searchall-minilm
Standard RAGnomic-embed-text
High qualitymxbai-embed-large
Document retrievalsnowflake-arctic-embed
Multilingualbge-m3

Pulling models in Ollama

ollama pull nomic-embed-text
ollama pull mxbai-embed-large

Generating embeddings

curl http://localhost:11434/api/embeddings -d '{
  "model": "nomic-embed-text",
  "prompt": "Docker is a container tool."
}'

Python example

import requests

url = "http://localhost:11434/api/embeddings"
payload = {
    "model": "nomic-embed-text",
    "prompt": "What is Ollama?"
}
response = requests.post(url, json=payload)
print(len(response.json()["embedding"]))

Token limits matter

Embedding models have token limits:

  • nomic-embed-text: 2048 tokens.
  • mxbai-embed-large: 512 tokens.
  • bge-m3: 8192 tokens.

Longer texts must be split into chunks.

Comparing quality

You can test quality yourself:

  1. Generate embeddings for the same texts using different models.
  2. Check similarity to relevant documents.
  3. Evaluate relevance of top results.
import numpy as np

def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

Memory and speed

ModelSpeedRAM/VRAM
all-minilmvery fastvery low
nomic-embed-textfastlow
mxbai-embed-largemoderatemoderate
bge-m3slowerhigher

Tips

  • For RAG with Ollama: nomic-embed-text or mxbai-embed-large.
  • Adjust chunk size to model token limit.
  • Use the same model for queries and documents.
  • Benchmark on your own data.
  • Vector dimension size affects memory usage in vector databases.

Further reading and resources

FAQ: Ollama embedding models

Which is the best embedding model? For most cases, nomic-embed-text. For highest quality, mxbai-embed-large.

Do I need a specialized model for RAG? Yes, the embedding model and LLM are separate components.

How many dimensions make sense? 384-1024 is sufficient for most applications.

Can I combine embeddings? Not directly. Only models with identical dimensions and similar vector space can be mixed.

What chunk size should I use? Depends on the model’s token limit, typically 256-1024 tokens.

Sources and further reading

Summary: Comparing Ollama embedding models

Ollama offers several embedding models for different requirements. nomic-embed-text is a solid all-rounder, mxbai-embed-large delivers higher quality, and all-minilm is extremely fast and small. Your choice depends on RAG quality needs, available hardware, and speed requirements. What matters most is chunk size, vector dimension, and using the same model consistently for both documents and queries. Testing a few models on your own data will quickly reveal the best fit.

Back to Blog
Share:

Related Posts