Skip to content
BotServBotServ
RAGRerankingCross-EncoderBi-EncoderRetrievalLocal AIQuality

Reranking: Improve RAG Results

What is reranking in RAG? How cross-encoders and rerankers enhance answer quality. Practical guide with examples and local models.

S

schutzgeist

11 min read
Reranking: Improve RAG Results

Reranking: Improving RAG Results

What This Article Covers

  • What reranking is and why it significantly improves RAG response quality
  • How cross-encoders and bi-encoders work and where they differ
  • Which reranking models suit local deployment
  • How to implement reranking with sentence-transformers in Python
  • When the extra effort pays off and what pitfalls to watch for

Introduction: Understanding Reranking

You’ve built a local RAG system, populated a vector database with your documents, and ask your first question. The answer is okay, but not really good. Often the problem isn’t the language model itself, but the text passages it receives as context. This is where reranking comes in.

Reranking is a second filtering step after the initial search. It takes the results from that first search and reorders them using a more precise model. The payoff: the truly relevant text passages rise to the top, and your language model gets better material to work with.

Why Do You Need Reranking?

Imagine asking your RAG system: “How do I enable automatic backups in application X?” The vector search returns 10 text passages. The relevant passage with the exact instructions sits at position 7 because the semantic similarity score wasn’t as high there. Passages 1 through 6 discuss “backup” and “security” in general terms, but not the specific feature you need.

Standard RAG takes the top 3 to 5 passages and sends them to the language model. The important passage at position 7 gets dropped. Your answer becomes vague or wrong.

With reranking, something different happens. A second model evaluates all 10 passages again, this time considering both the question and the text together. That passage at position 7 moves to position 1. The language model gets exactly the information it needs and delivers a precise answer.

Reranking Explained Simply

Reranking works like a two-stage hiring process. In the first round, a quick recruiter screens hundreds of resumes and selects 20 candidates. This selection is fast but rough. In the second round, an experienced hiring manager conducts detailed interviews with the top 10 candidates. It takes longer, but it’s far more accurate.

In RAG, the bi-encoder handles the first round: it converts the question and documents into vectors and compares them quickly. The cross-encoder handles the second round: it reads the question and document together and assesses relevance much more precisely. The result is a sorted list where the best matches come first.

Who Should Use Reranking?

Reranking makes sense for anyone working with RAG who wants to improve answer quality without swapping out the language model. It’s particularly valuable in these situations:

  • You have many documents and vector search often returns good but not the best matches
  • Your questions are specific and require precise text passages
  • You use a small or medium language model that struggles with poor context
  • You want to systematically test your RAG system’s quality

For simple demos with few documents, reranking isn’t essential. Once real data volumes come into play, though, it makes a noticeable difference.

Key Reranking Terminology

TermDefinition
RerankingSecond sorting step that re-evaluates the results from initial search
Cross-EncoderModel that processes question and document together and outputs a relevance score
Bi-EncoderModel that embeds question and document separately and compares vectors
RetrievalSearch step that fetches matching documents from the database
Relevance ScoreNumerical value indicating how relevant a document is to a question
Top-KNumber of results retained after search or reranking
Re-RankSynonym for reranking, the act of re-sorting results
CohereProvider of a well-known hosted reranking model
Sentence-TransformersPython library that can run reranking models locally
PrecisionMeasure of how many retrieved results are actually relevant

How Retrieval Works Without Reranking

In standard retrieval, you first store all documents as vectors in a database. An embedding model handles this. When you ask a question, it too becomes a vector. The database then finds the vectors closest to your question vector, typically using cosine similarity.

This approach is fast and scales well to millions of documents. But it has a weakness: the question and document are processed separately. The embedding model never sees both texts at the same time, so it can’t capture subtle relationships. A sentence like “The feature is disabled” might have high semantic similarity to “How do I enable this feature?” even though it doesn’t answer the question.

The result is coarse sorting. Top results are often relevant, but not always the most relevant. For many applications that’s fine. When precision matters, though, this is where reranking steps in.

How Reranking Works

Reranking uses a cross-encoder. Unlike a bi-encoder, a cross-encoder processes the question and document in a single pass. The model sees both texts at once and can precisely evaluate their relationship. It outputs a single relevance score that describes how well the document answers the question.

The workflow looks like this:

  1. Vector search delivers the top-K results, say 20 documents
  2. The cross-encoder computes a relevance score for each document in combination with the question
  3. Documents are reordered by this score
  4. The top N documents, for example 5, go to the language model

This second step is slower because the cross-encoder processes each document individually alongside the question. Since it only operates on a small set, overall time stays reasonable. Typically, reranking 20 documents takes a few milliseconds on a GPU and under a second on a CPU.

Bi-Encoder vs. Cross-Encoder

PropertyBi-EncoderCross-Encoder
ProcessingQuestion and document separatelyQuestion and document together
SpeedVery fast, scales to millionsSlower, only practical for small sets
AccuracyCoarse, good for initial filteringPrecise, captures subtle relationships
StorageVectors precomputed, minimal overheadNo precomputation possible, live computation
RoleInitial search in vector databaseSecond ranking of top-K results
ScalabilityHighLimited by computation time

These two approaches don’t compete; they complement each other. The bi-encoder does the fast initial screening, the cross-encoder does the precise refinement.

Reranking Models for Local Deployment

For local AI, several freely available reranking models are at your disposal. The main options:

BGE-Reranker from BAAI comes in a family of sizes. bge-reranker-base is compact and fast, while bge-reranker-large offers higher accuracy. Both are multilingual and work well with German text.

ms-marco-MiniLM from Microsoft was trained on the MS MARCO dataset. cross-encoder/ms-marco-MiniLM-L-6-v2 is lightweight and delivers solid results for English content. For German text, a multilingual model is usually the better choice.

jina-reranker from Jina AI is also multilingual and available in several sizes. jina-reranker-v2-base-multilingual covers many languages and is optimized for efficiency.

All these models can be loaded and run locally via the sentence-transformers library. You can also serve models with Ollama, but sentence-transformers is the more direct path for reranking specifically.

Example: RAG With and Without Reranking

A concrete comparison shows the difference. You ask: “How do I set the timeout for API requests in library Y?”

Without reranking, vector search returns these top 5 results:

  1. “General library configuration” (Relevant, but broad)
  2. “Timeout in network protocols explained” (Thematically related, not specific)
  3. “Installing the library” (Wrong context)
  4. “Error handling in API requests” (Partially relevant)
  5. “Example project with library Y” (No timeout content)

The relevant section “Setting timeout with setTimeout(5000)” lands at position 8 and never reaches the language model. The answer becomes vague.

With reranking, the new top 5 list looks like this:

  1. “Setting timeout with setTimeout(5000)” (Now position 1)
  2. “General library configuration”
  3. “Error handling in API requests”
  4. “Timeout in network protocols explained”
  5. “Advanced timeout options”

The language model gets the correct section first and delivers a precise answer with the right code example.

Implementation with sentence-transformers

Implementing reranking with sentence-transformers is straightforward. You only need the library and a reranking model.

from sentence_transformers import CrossEncoder

# Load model
model = CrossEncoder('BAAI/bge-reranker-base')

# Question and found documents
query = "How do I set the timeout for API requests?"
documents = [
    "General library configuration.",
    "Setting timeout with setTimeout(5000).",
    "Installing the library.",
    "Error handling in API requests.",
    "Example project with library Y."
]

# Create pairs of question and document
pairs = [[query, doc] for doc in documents]

# Calculate relevance scores
scores = model.predict(pairs)

# Sort results by score
ranked = sorted(zip(scores, documents), reverse=True)

for score, doc in ranked:
    print(f"{score:.4f}  {doc}")

The model returns a score for each pair. Higher values mean higher relevance. Then you sort the documents by score and keep the top N for the language model.

In a complete RAG pipeline, the flow looks like this: vector search returns 20 results, the cross-encoder reranks them, and the top 5 go to the language model. More on pipeline structure in RAG Basics.

When Is Reranking Worth It?

Reranking isn’t always necessary. This decision guide shows when it pays off:

  • Many documents: Once you have thousands of documents, initial retrieval becomes less precise, and reranking helps
  • Specific queries: The more targeted the question, the more it benefits from precise hits
  • Small language model: Models with small context windows need few but highly relevant sections
  • High quality demands: For support or legal topics, precision matters more than speed
  • Hybrid search already in place: If you already use hybrid search, reranking complements the combination of semantic and keyword search

Reranking makes less sense with very small document collections where vector search already returns good results, or in applications where latency is critical and every millisecond counts.

Common Reranking Pitfalls

  1. Wrong model for your language: An English-trained reranking model performs poorly on German text. Choose a multilingual model like BGE-Reranker or jina-reranker.

  2. Too many documents passed to reranking: Feeding 100 or more documents to the cross-encoder becomes slow without significant quality gains. 20 to 50 documents is a good compromise.

  3. Top-K before reranking too small: If vector search returns only 5 results and the relevant document is at position 6, reranking can’t save it. Fetch enough results first, say 20, before reranking.

  4. Top-K after reranking too large: If you pass 15 documents to the language model after reranking, less relevant sections dilute the answer. Keep only the top 3 to 5.

  5. Long documents not chunked: A cross-encoder has a maximum input length. If a document is too long, it gets truncated and the important part is lost. Chunking remains essential.

  6. No benchmark without reranking: Always compare results with and without reranking. That’s the only way to see if the extra step actually helps your data.

  7. Reranking model not on GPU: The cross-encoder is significantly slower on CPU. For frequent queries, a GPU is worth it.

  8. Misinterpreting scores: The absolute scores from a cross-encoder aren’t comparable across models. Use them only for sorting within a single model.

Hardware, Cost, and Security for Reranking

Reranking models are smaller than language models but larger than pure embedding models. bge-reranker-base has roughly 280 million parameters and runs on modern CPUs in acceptable time. For production systems with high query volume, a GPU is recommended.

Costs are minimal since all mentioned models are freely available. You only pay for hardware, which is an advantage of local AI. Compared to hosted reranking APIs like Cohere, you save ongoing fees but trade convenience for operational responsibility.

From a security standpoint, reranking is safe when run locally. Neither your questions nor documents leave your system. This is particularly important for sensitive data. Hosted reranking services transmit your text to external servers, which can be a risk with internal documents.

Further Reading and Resources on Reranking

FAQ: Reranking - Common Questions

What is reranking in RAG?

Reranking is a second sorting step that happens after vector search. A cross-encoder scores the question and document together, then reorders the results so the most relevant ones appear first.

Do I need reranking for small document collections?

With a small number of documents, vector search often produces good results on its own. Reranking becomes worthwhile once you have several hundred to thousands of documents, or when quality requirements are high.

What’s the difference between a bi-encoder and a cross-encoder?

A bi-encoder processes the question and document separately and is fast. A cross-encoder processes both together and is more accurate, but slower. Both complement each other in a RAG pipeline.

Which reranking model works well for German text?

Multilingual models like BGE-Reranker or jina-reranker-v2-base-multilingual work well for German content. English-only models such as ms-marco-MiniLM are less suitable for German.

How many documents should I rerank?

A good range is 20 to 50 documents. Too few risks missing the relevant document entirely. Too many slows down the process without proportional benefit.

Can reranking run on CPU?

Yes, most reranking models run on CPU. For production systems handling many requests, a GPU is significantly faster.

Can I use reranking with Ollama?

Ollama focuses on language and embedding models. For reranking specifically, sentence-transformers with a cross-encoder is the more direct approach. You can run both tools in parallel within a pipeline.

How much does reranking actually help?

It depends on your data. In many cases, reranking substantially improves precision on top results. Compare results with and without reranking using a test set to measure the effect.

Is reranking the same as hybrid search?

No. Hybrid search combines semantic and keyword search in the first stage. Reranking reorders the results afterward. Both approaches can be combined and reinforce each other.

Do I need to train a reranking model?

In most cases, no. Pretrained models like BGE-Reranker cover many use cases. For very specialized domains, fine-tuning can help, but it’s not necessary to start.

Does reranking noticeably increase response time?

Reranking 20 documents takes a few milliseconds on a GPU. On CPU it can take under a second. Compared to text generation by the language model, this overhead is usually small.

Sources and Further Reading

  • Sentence Transformers documentation on cross-encoders
  • BAAI BGE-Reranker on Hugging Face
  • Jina AI Reranker models
  • Microsoft MS MARCO cross-encoder models
  • Cohere Rerank API documentation for comparison with local models
Back to Blog
Share:

Related Posts