Embeddings with Ollama
What this article covers
- What embeddings are and why they matter.
- Which embedding models Ollama offers.
- How to generate embeddings via the API.
- Using embeddings in Python.
- Applications like RAG and semantic search.
Introduction: Embeddings with Ollama
Embeddings are numerical vectors that represent text in multidimensional space. Semantically similar content clusters together in this space. This property makes embeddings essential for RAG, semantic search, document clustering, and recommendation systems. Ollama makes it straightforward to run embedding models locally without relying on cloud services.
This article walks through which embedding models Ollama provides, how the API works, and how to use these vectors in your own projects.
Key terms
- Embedding: A vector representing text or other data.
- Embedding model: A specialized model for generating vectors.
- Vector space: A mathematical space where similarity is measured.
- Cosine similarity: A metric for proximity between two vectors.
- RAG: Retrieval-Augmented Generation, incorporating external documents into prompts.
- Vector database: Storage for embeddings.
Embedding models in Ollama
Ollama provides several embedding models. The most popular ones are:
- nomic-embed-text: 137 million parameters, solid performance, freely usable.
- mxbai-embed-large: More powerful multilingual model.
- bge-m3: Multilingual, suited for longer documents.
Download models the same way you would language models:
ollama pull nomic-embed-text
ollama pull mxbai-embed-large
Generating embeddings
Via REST API
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": "Ollama is a tool for local AI."
}'
Response:
{
"model": "nomic-embed-text",
"embeddings": [[0.12, -0.34, 0.56, ...]]
}
Multiple texts
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": [
"First text.",
"Second text.",
"Third text."
]
}'
Python example
import requests
def embed(texts, model="nomic-embed-text"):
url = "http://localhost:11434/api/embed"
payload = {"model": model, "input": texts}
resp = requests.post(url, json=payload)
return resp.json()["embeddings"]
vectors = embed(["Hello world", "Using Ollama locally"])
print(len(vectors), len(vectors[0]))
Computing similarity
Use numpy to calculate cosine similarity between two vectors:
import numpy as np
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
v1 = np.array(vectors[0])
v2 = np.array(vectors[1])
print(cosine_similarity(v1, v2))
Ollama local vs OpenAI
| Feature | Ollama local | OpenAI API |
|---|---|---|
| Cost | Electricity only | Per token |
| Privacy | Stays local | Sent to cloud |
| Speed | Hardware dependent | Fast |
| Quality | Good | Very good |
| Scaling | Limited by your hardware | Highly scalable |
For many use cases, local Ollama embeddings are sufficient. If you need maximum accuracy or process very large datasets, you can supplement with cloud models.
Applications
RAG
In a RAG system, documents are split into chunks, converted to embeddings, and stored in a vector database. When you ask a question, it gets converted to a vector too, and the most relevant chunks are retrieved.
Semantic search
Instead of keyword matching, you find semantically similar content. This works across languages as well.
Classification and clustering
Embeddings let you categorize texts or automatically group them by topic.
Vector databases
Combine Ollama embeddings with local vector databases:
- ChromaDB: Easy to use, popular in Python.
- Qdrant: Fast, scalable vector database.
- Weaviate: Model-agnostic, strong for enterprise.
- pgvector: Vector extension for PostgreSQL.
Tips
- Break text into meaningful chunks.
- Adjust chunk size and overlap based on your use case.
- Choose an embedding model suited to your language.
- Store embeddings once, don’t regenerate them for each query.
- Measure recall and precision with sample queries.
Common pitfalls
- Wrong model: Not every model is suitable for embeddings.
- Text too long: Longer text must be chunked beforehand.
- Varying dimensions: Different models produce vectors of different lengths.
- No normalization: Cosine similarity requires normalized vectors.
- Wrong API endpoint:
/api/embedis the Ollama endpoint, not/v1/embeddings.
Further reading and resources
- BotServ.de Ollama REST API
- BotServ.de Ollama commands
- BotServ.de RAG knowledge base
- BotServ.de Vector databases
FAQ: Embeddings with Ollama
Which embedding model should I use?
nomic-embed-text is a solid starting point. mxbai-embed-large is more powerful but larger.
Can I embed entire documents at once? It’s better to chunk them first.
Are Ollama embeddings GDPR compliant? Yes, since everything is processed locally.
How do I store embeddings? In a vector database like ChromaDB, Qdrant, or pgvector.
Which similarity metric should I use? Cosine similarity is the standard choice.
Sources and further reading
- Ollama API Docs: https://github.com/ollama/ollama/blob/main/docs/api.md
- nomic-embed-text: https://ollama.com/library/nomic-embed-text
- Sentence Transformers: https://sbert.net/
Summary: Embeddings with Ollama
Ollama lets you run embedding models locally and convert text into vectors. This is the foundation for RAG, semantic search, and text classification. nomic-embed-text is a popular entry point. When you split text into appropriate chunks, store vectors in a local vector database, and measure similarity with cosine distance, you can quickly build a capable local RAG system.


