Skip to content
BotServBotServ
RAGHybrid SearchBM25Dense RetrievalKeyword SearchVectorLocal AI

Hybrid Search: Combining Semantic and Lexical

What is Hybrid Search in RAG? Combining vector and keyword search for better results. BM25, Dense Retrieval explained.

S

schutzgeist

12 min read
Hybrid Search: Combining Semantic and Lexical

Hybrid Search: Combining Semantic and Lexical Methods

What this article covers

  • Understanding what hybrid search is and why it matters for RAG systems
  • The difference between lexical and semantic search approaches
  • How to combine both methods, including techniques like Reciprocal Rank Fusion
  • Concrete code examples using Chroma, Qdrant, and LanceDB
  • Common pitfalls and how to avoid them

Introduction: Hybrid Search Explained

In local RAG, the goal is to locate relevant passages from your own documents and provide them as context to a language model. The quality of answers depends heavily on how well the search performs. Hybrid search combines two search strategies, each with distinct strengths and limitations: lexical search and semantic search. Together, they deliver significantly better results than either method alone.

For a deeper foundation, see RAG Fundamentals and the overview article What is local AI?.

Imagine searching through technical documentation for the term “Error 404”. Pure vector search, also called semantic search, grasps the meaning of your question. It finds sections about “page not found” or “page unavailable” because the concept is semantically related. But it might miss the exact error code 404, since a number carries little semantic weight in vector space.

Pure keyword search, also called lexical search, finds every document containing the word “404”. It reliably locates the error code. Yet it misses all sections that describe the problem without mentioning the number. A section titled “Page not found” won’t appear if it lacks the string “404”.

Hybrid search runs both queries and combines the results. This way, you find both the exact error code and the semantically related explanations. This is especially valuable for technical documentation, API references, and manuals where codes, names, and technical terms coexist with descriptive prose.

Hybrid Search at a Glance

Hybrid search works like a librarian searching in two ways simultaneously. First, he asks: “Which books cover the same subject as your question?” That’s semantic search, focused on meaning and context. Second, he asks: “Which books contain exactly the words you mentioned?” That’s lexical search, focused on exact matches.

Both result sets are merged and ranked by relevance. A document that ranks high in both lists rises to the top of the combined results. A document that performs well in only one list gets lower weight. This way, search benefits from both strengths at once.

Hybrid search is for anyone building RAG systems who wants to improve search quality. This includes developers creating internal knowledge bases, teams making documentation searchable, and users setting up local AI solutions for proprietary documents. Use cases mixing descriptive text with structured information, like codes, IDs, names, and technical terms, benefit the most.

If you’re new to RAG, start with RAG Fundamentals and Embedding Models before diving into hybrid search.

TermDefinition
Hybrid SearchCombination of lexical and semantic search
Dense RetrievalSemantic search over dense vectors, also called embeddings
Sparse RetrievalLexical search over sparse vectors with only certain dimensions populated
BM25Enhanced TF-IDF algorithm for keyword search
TF-IDFMeasure of word importance in a document
EmbeddingNumerical vector representing text meaning
KeywordSearch term that must appear exactly in the text
SemanticBased on meaning, not exact word matching
FusionMethod for combining multiple result lists into one
RRFReciprocal Rank Fusion, a common fusion method

Lexical search is classical text search. It looks for exact word matches between query and document. The best-known algorithms are TF-IDF and BM25.

TF-IDF stands for Term Frequency, Inverse Document Frequency. The algorithm scores how often a word appears in a document and weights it higher if it occurs in few documents overall. A word appearing everywhere carries less signal.

BM25 improves upon TF-IDF. It accounts for document length and dampens the effect of very common words, leading to more relevant hits. BM25 remains the standard for keyword search today and powers search engines like Elasticsearch and Lucene.

Lexical search shines with exact matches. It reliably finds names, error codes, version numbers, IDs, and technical terms. If a document doesn’t contain the search term, it won’t appear in results. For many queries, that’s the desired behavior.

Semantic search works with embeddings. An embedding model converts text into a vector representing its meaning. For a search query, a vector is also created, and the vector database finds the nearest vectors. Learn more in Embedding Models.

Semantic search’s strength lies in understanding meaning. It recognizes synonyms, paraphrases, and topical relationships. A query like “How do I stop the server?” finds sections about “shut down the server” or “kill the process”, even though the words differ.

Semantic search falters with specific terms. An error code like “404” or a product name like “Model-X100” carries little semantic meaning. The vector for a number or ID barely differs from other numbers or IDs. Semantic search finds these terms poorly or not at all.

Weaknesses of Each Approach

Neither method is optimal on its own. Each has blind spots.

Lexical search fails when the query and document use different words for the same concept. A search for “auto” won’t find a document containing only “vehicle”. A search for “server stop” won’t find a document about “kill process”. It understands neither synonyms nor paraphrases.

Semantic search fails with exact terms carrying no semantic meaning. Error codes, version numbers, IDs, product names, and abbreviations are hard to distinguish in vector space. A search for “HTTP 500” might return all sections on server errors, but not the one containing exactly that code.

Hybrid search solves this by running both searches and combining results. Each method’s strengths cover the other’s weaknesses.

How Hybrid Search Works

Hybrid Search operates in three steps:

  1. Lexical Search: The query is matched against all documents using BM25 or a similar algorithm. The result is a ranked list based on keyword relevance.
  2. Semantic Search: The query is converted into a vector and matched against all embeddings in the vector database. The result is a ranked list based on vector similarity.
  3. Fusion: Both ranked lists are combined into a single list. Documents that rank well in both lists appear near the top.

Fusion is the critical step. Several methods exist for combining two ranked lists. The two most common are weighted sum and Reciprocal Rank Fusion.

With weighted sum, each list receives a weighting factor. For example, 0.7 for semantic search and 0.3 for lexical search. The scores are added together to produce a final score. The problem: the scores from both searches often operate on different scales, making the weighting difficult to calibrate.

Reciprocal Rank Fusion solves this by combining ranks rather than scores. Scale becomes irrelevant.

Reciprocal Rank Fusion (RRF)

RRF is the most commonly used fusion method for Hybrid Search. It combines ranked lists without needing the actual scores. This makes it robust and straightforward to apply.

The formula is:

RRF(d) = sum over all ranked lists of: 1 / (k + rank(d))

Here, d is the document, rank(d) is the document’s position in the respective ranked list, and k is a constant, typically 60. A document ranking first in both lists receives a high RRF score. A document ranking first in only one list receives a moderate score.

The parameter k controls how strongly top ranks are weighted. A small k makes top positions more dominant, while a large k distributes influence more evenly. The value 60 has proven effective in practice and is used as a standard in many systems.

RRF is popular because it requires no score normalization. BM25 scores and cosine similarities operate on completely different scales. RRF ignores the scores and uses only the ranks. This makes it robust against scaling issues.

Implementation with Local Tools

Several local vector databases support Hybrid Search natively. Below you’ll find examples for Chroma, Qdrant, and LanceDB. More about these databases in Vector Databases.

Chroma

Chroma offers Hybrid Search by combining embedding search with its own BM25 implementation. More details in Chroma.

import chromadb

client = chromadb.PersistentClient(path="./chroma_db")
collection = client.get_or_create_collection("dokumente")

# Dokumente hinzufuegen
collection.add(
    documents=["Error 404: Seite nicht gefunden", "Server stoppen und neu starten"],
    metadatas=[{"quelle": "docs"}, {"quelle": "docs"}],
    ids=["doc1", "doc2"]
)

# Semantische Suche
results_semantic = collection.query(
    query_texts=["Seite nicht gefunden"],
    n_results=5
)

# Lexikalische Suche ueber where-Filter oder externes BM25
# Chroma unterstuetzt Hybrid Search ueber Custom Embedding Functions
# oder externe BM25-Implementierungen wie rank_bm25
from rank_bm25 import BM25Okapi

docs = collection.get()["documents"]
tokenized = [doc.lower().split() for doc in docs]
bm25 = BM25Okapi(tokenized)
scores = bm25.get_scores("404 seite nicht gefunden".split())

In practice, you combine the results of both searches using RRF. Chroma increasingly offers native Hybrid Search support, but combining results manually with rank_bm25 remains a proven approach.

Qdrant

Qdrant supports Hybrid Search natively through sparse and dense vectors. More in Qdrant.

from qdrant_client import QdrantClient
from qdrant_client.models import SparseVector, SearchRequest, FusionQuery

client = QdrantClient(path="./qdrant_db")

# Hybride Suche mit nativer Fusion
results = client.query_points(
    collection_name="dokumente",
    prefetch=[
        SearchRequest(
            using="dense",
            vector=[0.1, 0.2, 0.3],  # Embedding der Suchanfrage
            limit=20
        ),
        SearchRequest(
            using="sparse",
            vector=SparseVector(
                indices=[10, 25, 80],
                values=[0.8, 0.5, 0.3]
            ),
            limit=20
        )
    ],
    query=FusionQuery(fusion="rrf"),
    limit=10
)

Qdrant performs the fusion directly on the server. You don’t need to implement RRF yourself. This is especially convenient when working with large document collections where fusion efficiency matters.

LanceDB

LanceDB also offers Hybrid Search with integrated RRF support.

import lancedb

db = lancedb.connect("./lancedb_db")
table = db.open_table("dokumente")

# Hybride Suche mit FTS und Vektor-Suche
results = table.search(
    query="Error 404",
    query_type="hybrid"
).limit(10).to_list()

LanceDB combines Full-Text Search and vector search in a single call. Fusion happens automatically via RRF.

Example: Hybrid Search in Documentation

Imagine you have documentation with these sections:

  • Section A: “Error 404: The requested page was not found.”
  • Section B: “When a page doesn’t exist, the server returns an error.”
  • Section C: “Error 500: Internal server error during processing.”

The search query is: “What does Error 404 mean?”

Lexical Search (BM25):

  1. Section A, contains “Error” and “404” exactly
  2. Section C, contains “Error”, but not “404”
  3. Section B, contains neither word

Semantic Search (Vector Search):

  1. Section B, semantically closest, “page doesn’t exist” matches the question
  2. Section A, semantically related, contains the error
  3. Section C, semantically related, but a different error

Hybrid Search (RRF):

  1. Section A, well-positioned in both lists
  2. Section B, ranked first in the semantic list
  3. Section C, moderate ranking in both lists

Section A ends up at the top because it covers both the exact code and the semantic meaning. This is the advantage of Hybrid Search: the best result from both worlds ends up right at the top.

Hybrid Search is powerful, but there are several traps to watch for.

  1. Incorrect Weighting: If you use weighted sum instead of RRF, wrong weights can skew results. Test different weights with real search queries.
  2. Tokenization for BM25: German compound words like “Datenbankverbindung” aren’t split by simple tokenizers. A stemmer or German-specific tokenizer significantly improves keyword search.
  3. Mismatched Chunk Sizes: If lexical and semantic search operate on different chunk sizes, results become hard to compare. Both searches should use identical chunks.
  4. Missing Indexing: BM25 requires an inverted index. If you only have a vector database, you must build the index for keyword search separately.
  5. Language Dependency: BM25 is language-dependent. An English tokenizer performs poorly on German text. Make sure your tokenization matches your documents’ language.
  6. Too Many Results: If both searches each return 100 results, fusion becomes slow and unwieldy. Limit results per search to 20 to 50.
  7. No Reranking: Hybrid Search delivers good candidates, but the order isn’t perfect. A downstream reranking model improves quality further. More in Reranking.

Hybrid Search requires two indexes: a vector index for semantic search and an inverted index for lexical search. Together they consume more storage than a single index, but the overhead is manageable. BM25 indexes are compact and fast.

Compute: Semantic search needs an embedding model to vectorize the query. This runs efficiently on a CPU. Lexical search is very fast and demands minimal compute. The fusion step itself is trivial and adds negligible latency.

Costs: All required tools are open source and free. Chroma, Qdrant, and LanceDB are released under permissive licenses. BM25 implementations like rank_bm25 are freely available as well.

Security: When all components run locally, your documents stay on your machine. Neither queries nor documents leave your system. This is a core advantage of local AI, as explained in What is local AI?.

FAQ: Hybrid Search - Common Questions

What is hybrid search in simple terms?

Hybrid Search combines traditional keyword search with semantic vector search. Both searches run simultaneously and results are merged. This gives you the benefits of exact matches and semantic understanding at once.

Do I need hybrid search for every RAG system?

Not necessarily. For plain prose without code or proper names, semantic search alone often suffices. But if your documents contain error codes, IDs, product names, or other exact terms, hybrid search becomes recommended.

What’s the difference between dense and sparse retrieval?

Dense retrieval uses dense vectors, where embeddings have values across all dimensions. Sparse retrieval uses sparse vectors with only a few dimensions populated, typically for keyword search. Hybrid search combines both approaches.

What is RRF and why is it used so often?

RRF stands for Reciprocal Rank Fusion. It combines rank lists without needing the actual scores. This makes it robust across different score scales. It’s straightforward to implement and delivers good results in practice.

Do I have to implement BM25 myself?

No. Ready-made libraries like rank_bm25 exist for Python. Qdrant and LanceDB offer BM25 directly in the database. Chroma can be paired with external BM25 libraries.

How do I choose the weight between vector and keyword search?

With RRF, you don’t need to set weights; the method combines automatically. With weighted sum, start at 0.5 for each and test against real queries. Adjust the weights until results match your expectations.

Does hybrid search work with German text?

Yes, but tokenization for BM25 must be configured for German. A German stemmer and stopword list improve keyword search. For semantic search, use a multilingual embedding model.

Is hybrid search slower than a single search?

Yes, two searches run. But the speed difference is small since both can execute in parallel. The fusion itself is very fast. In practice, the difference is barely noticeable.

Should I use reranking after hybrid search?

Yes, it’s recommended. Hybrid search produces good candidates, but a reranking model can improve their order further. Reranking scores the top results again using a stronger model.

Can I use hybrid search without a GPU?

Yes. Lexical search needs no GPU. Semantic search requires an embedding model that runs on CPU. For small to medium document collections, CPU performance is sufficient.

Which vector databases natively support hybrid search?

Qdrant and LanceDB support hybrid search with integrated RRF out of the box. Chroma can be extended with external BM25 libraries. Weaviate and Milvus also offer native hybrid search features.

Sources and Further Reading

  • Original Reciprocal Rank Fusion paper by Cormack et al.
  • Qdrant documentation on hybrid search
  • LanceDB documentation on hybrid search
  • Chroma documentation
  • BM25 original publication by Robertson and Zaragoza
  • Sentence Transformers documentation
Back to Blog
Share:

Related Posts