Embedding Models for Local AI
What this article covers
- What embeddings are and how they work.
- How text, images, or mixed data get converted into vectors.
- Which embedding models work well for German-language content.
- How embeddings power RAG and semantic search.
- Model selection, dimensions, metrics, and common pitfalls.
Introduction: Embedding models for local AI
Embedding models are the backbone of many AI applications. They transform text, images, or other data into vectors, sequences of numbers that capture the meaning of content. Similar content lands close together in vector space, which enables search, clustering, and recommendations.
RAG systems, semantic search, and document analysis all rely on embeddings. A solid embedding model often matters more than a massive language model. Even the best LLM can only answer queries correctly if the retrieval step finds the right chunks in the first place.
Why do you need embedding models?
If you want to search documents, images, or audio clips, simple keyword matching won’t cut it. Embeddings enable semantic search. You can search for concepts instead of exact words. Consider these examples:
- “What is the notice period?” finds the relevant contract clause even if phrased differently.
- “GPU installation problems” finds the right troubleshooting section.
- “Similar customer complaints” locates earlier tickets with similar meaning.
Embeddings make unstructured data searchable.
Embedding models explained
An embedding model takes input text and returns a vector of fixed length, called its dimension. Common dimensions range from 384 to 4096. Most models use a Transformer backbone that encodes meaning as numbers.
Key terms:
- Embedding: A vector representing the meaning of content.
- Dimension: The length of the vector.
- Vector database: Storage that searches embeddings efficiently.
- Semantic search: Finding meaning rather than exact words.
- Cosine similarity: A measure of how close two vectors are.
- Reranking: Re-scoring results after retrieval to find the best ones.
Who should use embedding models?
- Developers building RAG systems.
- Teams needing to search documents or tickets.
- Researchers analyzing data.
- Anyone wanting to run semantic search locally.
Key concepts around embeddings
- Sentence Transformers: A library and model class for sentence-level embeddings.
- E5, BGE, GTE: Popular embedding model families.
- Multilingual: Support for multiple languages, including German.
- MTEB: A benchmark for evaluating embedding models.
- Sparse vectors: Vectors with few non-zero values, such as BM25.
- Dense vectors: Standard embeddings with many values.
Common embedding models
all-MiniLM-L6-v2
Small and fast. 384 dimensions. Suitable for prototypes and smaller datasets. German works, but not as well as specialized multilingual models.
BGE-m3
A multilingual model from BAAI. Strikes a good balance between quality and size. 1024 dimensions. Also supports sparse retrieval. Strong for RAG.
nomic-embed-text-v1.5
An open-source embedding model with solid performance. 768 dimensions. Good German quality and a permissive license.
e5-mistral-7b-instruct
A large model with very high quality. Requires significantly more resources. Only worthwhile if maximum quality is essential.
Practical example: Creating embeddings with Python
You can quickly generate embeddings using sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('BAAI/bge-m3')
texts = ['Hello world', 'Local AI respects data privacy']
embeddings = model.encode(texts)
print(embeddings.shape)
The output shows the vector dimension and number of texts.
Storing embeddings in vector databases
After creation, embeddings are stored in a vector database. When you search, the query is also converted to a vector. The database finds the nearest neighbors. Popular options include:
- Chroma: Simple, good for prototypes.
- Qdrant: Fast, scalable, Docker-ready.
- pgvector: An extension for PostgreSQL.
- Weaviate: Enterprise features, modular design.
Dimensions, model size, and speed
- Smaller dimensions: Less memory, faster search, lower quality.
- Larger dimensions: More memory, slower search, better quality.
- Quantization: Reduces vector size with minimal quality loss.
For most local RAG systems, 768 to 1024 dimensions is sufficient.
Common pitfalls with embedding models
- Wrong language: English models perform poorly on German text.
- Chunks too small: Context gets lost, embeddings become imprecise.
- Wrong metric: Cosine similarity is standard but not always optimal.
- No reranking: Top results are good, but reranking improves them further.
- Passage vs. query: Some models need different prompts for queries and documents.
- License: Not all models are suitable for commercial use.
Further reading and resources
- BotServ.de Local RAG
- BotServ.de Vector Databases
- BotServ.de Reranking
- Sentence Transformers
- MTEB Benchmark
FAQ: Embedding models
How do embeddings differ from an LLM? An LLM generates text; an embedding model generates vectors.
How large should an embedding model be? For most use cases, 768 to 1024 dimensions is plenty. Larger is rarely needed.
Can I create embeddings on a CPU? Yes, especially with smaller models.
Which model family do you recommend? BGE, E5, GTE, and nomic-embed-text are all solid starting points.
Are embeddings possible for images? Yes, using multimodal models like CLIP.
How long does it take to index many documents? Depending on the model and hardware, anywhere from minutes to hours. Much faster with a GPU.
Sources and further reading
- Sentence Transformers: https://www.sbert.net/
- MTEB: https://huggingface.co/spaces/mteb/leaderboard
- BGE: https://huggingface.co/BAAI
- nomic-embed-text: https://huggingface.co/nomic-ai/nomic-embed-text-v1.5
Summary: Embedding models for local AI
Embedding models form the foundation of semantic search and RAG. They transform text into vectors and enable meaning-based retrieval. Multilingual models like BGE, nomic-embed-text, and E5 work well for German-language applications. Language choice, dimensionality, chunking strategy, and the right vector database all matter. Select the right model, and you’ve laid the groundwork for effective RAG results.


