Skip to content
BotServBotServ
RAGQualityTestingMetricsRecallPrecisionFaithfulnessRAGASLocal AI

Testing RAG Quality: Metrics and Methods

How to test your RAG system quality? Learn Recall, Precision, Faithfulness, Answer Relevance and practical testing methods.

S

schutzgeist

14 min read
Testing RAG Quality: Metrics and Methods

Testing RAG Quality: Metrics and Methods

What this article covers

  • Learn why quality testing is essential for RAG systems and what happens when you skip it
  • Understand key metrics like Recall, Precision, Faithfulness, and Answer Relevance through concrete examples
  • Discover how the RAGAS framework works and how to use it to evaluate your system
  • Get a step-by-step guide to building and assessing your own test set
  • Identify common pitfalls and how to avoid them

Introduction: Understanding RAG quality testing

You’ve built a RAG system, perhaps with Ollama and a local vector database. On the surface, everything works. You ask a question, get an answer that sounds reasonable. But how good is it really? How often does it give wrong answers? And where do the failures come from, in the retrieval step or in answer generation?

This is where quality testing comes in. In this article, I’ll walk you through how to systematically measure and improve your RAG system’s quality. We’ll look at both the metrics and their practical implementation. If you’re not yet sure what RAG actually is, start with the RAG fundamentals article first.

Why do you need quality tests?

Imagine you’re running a RAG system for internal support at your company. Employees ask questions about vacation requests, expense policies, and IT problems. The system answers 80 percent correctly. The other 20 percent are wrong or incomplete.

Without systematic testing, you don’t know which 20 percent those are. An employee asks about the deadline for vacation requests and gets the answer “three weeks before start date”. It sounds plausible, but the correct answer is “four weeks”. Nobody complains because the response is convincing. The error only surfaces when someone submits too late and their request is rejected.

Quality tests prevent this. They show you where your system is weak before it causes problems in production. You spot error patterns and can make targeted improvements, like expanding your knowledge base in a certain area or optimizing reranking.

RAG quality testing explained simply

Think of a student preparing for an exam. They could just claim they’re ready. Or they practice with old test papers and count how many questions they answer correctly. That’s the only way to know if they’re really prepared.

A quality test for RAG works the same way. You gather a collection of questions where you know the correct answers, feed them to your system, and compare the output against the expected answer. From this comparison, you calculate metrics that tell you how well the system performs. A single number isn’t enough because RAG involves multiple steps. Retrieval can be good and generation bad, or vice versa. That’s why you measure both steps separately.

Who are quality tests for?

Quality tests serve several audiences:

  • Developers building a RAG system who want to know if their configuration works well
  • Operators of small and medium systems who can’t afford expensive evaluation teams but still need confidence that their system works reliably
  • Teams migrating from an API-based solution to local AI and wanting to compare quality
  • Beginners building a RAG system for the first time who need to understand what quality looks like

You don’t need a PhD in statistics. The core ideas are straightforward, and tools like RAGAS do much of the heavy lifting for you.

Key terms in RAG quality

TermMeaning
RecallThe proportion of relevant passages your system found. High Recall means it finds nearly everything important.
PrecisionThe proportion of found passages that are actually relevant. High Precision means little unnecessary noise.
F1Combines Recall and Precision into one score. Useful when you need a single number for comparison.
FaithfulnessMeasures whether the answer contains only statements backed by the source passages. No hallucinations.
Answer RelevanceMeasures whether the answer actually addresses the question asked, rather than just saying something relevant.
Context PrecisionEvaluates whether the retrieved passages are useful for answering the question.
Context RecallEvaluates whether all necessary information from the sources was found.
RAGASA framework for automatic RAG system evaluation that computes multiple metrics.
Ground TruthThe correct reference answer you use to test your system against.
Evaluation SetA collection of test questions with corresponding reference answers and expected sources.

Why RAG quality is hard to measure

RAG has two main steps: retrieval and generation. Each can be good or bad, and the combination makes evaluation complex. A system can find perfect passages and still generate a wrong answer. Or it retrieves mediocre sources but formulates a usable answer from them.

Quality is also subjective. What counts as a “good answer” depends on context. A brief answer might be ideal for a quick question but inadequate for a complex one. Metrics can only partially capture this subjectivity.

Another challenge is distinguishing between hallucination and wrong retrieval. If an answer is incorrect, is it because the model invented something or because the wrong passages were retrieved? Without measuring both steps separately, you stay in the dark. That’s why you use multiple metrics at once, each illuminating different aspects.

Retrieval metrics

The retrieval phase is the first step: your system searches the vector database for passages matching the question. Two metrics are central here.

Recall: Did we find the right passages?

Recall measures how many relevant passages your system found. Suppose your database has five passages important for a question. Your system finds three. Recall is 60 percent.

High Recall matters because generation can only work with what was found. If a critical passage is missing, the model can’t use it, and your answer becomes incomplete or wrong.

Precision: Are the found passages relevant?

Precision measures how many of the retrieved passages are actually relevant. Your system finds ten passages, but only four matter for the question. Precision is 40 percent.

Low Precision means the model must work with many irrelevant passages. This can degrade the answer because the model gets confused or loses focus. It also wastes compute and time processing unnecessary passages.

Methods like Hybrid Search and reranking help improve Precision without lowering Recall. Hybrid Search combines semantic and keyword-based search, while reranking re-orders retrieved passages by relevance.

Generation Metrics

Once the retriever finds relevant passages, the language model generates an answer based on them. Two key metrics apply here as well.

Faithfulness: Does the answer match the sources?

Faithfulness measures whether every claim in the answer is supported by the retrieved passages. If the answer states that vacation request deadlines are three weeks, but the source says four weeks, the answer lacks faithfulness. The model either misread the source or added information on its own.

Low faithfulness signals hallucination. The model invents facts not present in the sources. This is particularly risky because such answers often sound convincing. Source citations in the output can make this problem more transparent to users, but they don’t solve it.

Answer Relevance: Does the answer actually address the question?

Answer Relevance measures whether the answer truly responds to the user’s question. An answer can be fully faithful yet still miss the mark. If someone asks “How long is the notice period?” and the answer explains the termination process in detail but never states the actual notice period, then the answer isn’t relevant.

Answer Relevance is harder to measure than Faithfulness because “relevance” is more subjective. RAGAS uses an approach where new questions are generated from the answer, then compared to the original question. A strong answer should make it possible to reconstruct the original question from it.

RAGAS Framework

RAGAS (Retrieval Augmented Generation Assessment) is an open-source framework that automates evaluation of RAG systems. It calculates the four most important metrics in a single run:

  • Faithfulness: Are all claims in the answer supported by the context?
  • Answer Relevance: Does the answer address the question?
  • Context Precision: Are the retrieved passages useful for answering?
  • Context Recall: Were all necessary pieces of information found?

How RAGAS works

RAGAS uses a language model as the evaluator. You provide the question, the retrieved passages, the generated answer, and optionally a reference answer. The evaluation model analyzes these inputs and produces scores for each metric, typically between 0 and 1. RAGAS itself therefore needs a language model. With local AI, you can use a local model as the evaluator, for instance via Ollama. This consumes compute resources but keeps all data on your system.

Using RAGAS

Typical usage works like this: you create an evaluation set, which is a list of questions with reference answers. For each question, you run your RAG system and store the retrieved passages and the generated answer. You pass this data to RAGAS, which computes the metrics. At the end you get a report with average scores and individual results per question.

RAGAS is available in Python and integrates into existing pipelines. For smaller systems, running evaluation manually as needed is sufficient. For larger systems, automate the process and run it regularly, for example after each knowledge base update.

Creating a Test Set

A good test set is the foundation of any evaluation. Without meaningful test questions, you measure nothing useful. Here are the steps for building one.

Step 1: Gather questions

Collect 20 to 50 questions that are typical for how your system will be used. Draw from multiple sources: actual user questions if you have logs, your own test questions, questions from different topic areas, and a mix of simple and complex ones. Make sure your questions cover the breadth of what the system needs to handle, not just easy cases.

Step 2: Write reference answers

For each question, write the best possible answer based on your understanding of the sources. The reference answer doesn’t need to be a word-for-word copy from your knowledge base, but it must be correct and complete. This reference is your ground truth.

Step 3: Document expected sources

For each question, note which passages from your knowledge base should be relevant. This helps later when evaluating retrieval performance. If the system finds different passages, you can tell whether the search needs improvement.

Step 4: Include edge cases

Deliberately add difficult questions: ones where the answer requires combining multiple passages, questions with ambiguous terms, questions where no answer exists in the knowledge base. Edge cases without answers are especially valuable because that’s where the system should ideally say “I don’t have information on that” rather than invent something.

Example: Evaluating a RAG System

Here’s a concrete workflow for evaluation with RAGAS.

Setup

You’ve built your RAG system using Local RAG. Your test set contains 30 questions with reference answers and expected sources. RAGAS is installed and you’ve configured a local evaluation model via Ollama.

Execution

For each question in your test set, query your RAG system, save the retrieved passages (Context) and generated answer (Answer), and note the reference answer (Ground Truth). Collect this data in a table or JSON file.

Evaluation with RAGAS

Pass the collected data to RAGAS. The framework calculates all four metrics for each question and gives you averages. A typical result might look like this:

  • Faithfulness: 0.82
  • Answer Relevance: 0.75
  • Context Precision: 0.68
  • Context Recall: 0.71

Interpretation

The numbers reveal weaknesses in retrieval. Context Precision and Context Recall are relatively low. The system doesn’t consistently find the best passages. Faithfulness is higher, meaning the model generally reproduces the found passages correctly. Answer Relevance falls in between, suggesting some answers don’t optimally address the question, possibly because the sources weren’t ideal.

The takeaway: focus on improving retrieval first, perhaps through reranking or hybrid search, before refining generation.

Manual Testing

Not everyone needs a framework like RAGAS right away. Manual testing is a good starting point and often more revealing than expected.

Workflow

  1. Take 20 questions from your test set
  2. Query your RAG system with each one
  3. Compare the answer to your reference answer
  4. Mark each as: correct, partially correct, or wrong
  5. For wrong answers, note the cause (bad retrieval, bad generation, question misunderstood)

Analysis

After 20 questions you have simple statistics. Maybe 14 correct, 4 partially correct, 2 wrong. More important than the percentage is the error analysis. Look closely at the wrong and partially correct answers. Do you see patterns? Perhaps all wrong answers fail on questions about a specific topic, pointing to a gap in the knowledge base. Or the retrieved passages were right but the answer was poorly formulated, indicating a generation problem.

Manual testing takes time, but it gives you a feel for the system that no metric can replace. For an initial assessment, it’s more than enough.

Common Pitfalls in Quality Testing

Quality testing has several pitfalls that can skew your results.

Test sets that are too small: With five questions, you won’t get meaningful metrics. A single outlier dramatically changes the outcome. Use at least 20 questions, ideally 50 or more.

Unrealistic questions: If your test questions are all straightforward, the system looks better than it actually is. Include difficult and ambiguous questions that really challenge the model.

Missing edge cases: Questions with no answer in your knowledge base are often overlooked. The system should recognize when it lacks information instead of hallucinating. Test this explicitly.

Outdated reference answers: When your knowledge base changes, reference answers can become incorrect. Update your test set regularly, especially after major updates.

Weak evaluation model: Using RAGAS with a small local model as your evaluator can produce unreliable scores. Test with different evaluation models and compare the results.

Overfitting to metrics: If you only watch the numbers and optimize your system to improve metrics, actual quality can suffer. Metrics are tools, not the goal. Always include manual spot checks.

One-time testing: A single test after building your system isn’t enough. Any change to the system, knowledge base, or parameters can affect quality. Run tests regularly, ideally after every significant change.

Hardware, Costs, and Security in Quality Testing

Quality testing consumes resources, especially when you use RAGAS with a local evaluation model. Each question in your test set requires multiple calls to the evaluation model, since RAGAS performs separate analyses for each metric. With 50 questions and four metrics, you’re looking at 200 model calls per evaluation.

For hardware, this means you need enough RAM and VRAM to run the evaluation model smoothly. A model with 7 to 8 billion parameters, like Llama 3 8B, is a good compromise between quality and resource usage. It runs on a system with 16 GB RAM, though not particularly fast. For comfortable speed, 32 GB RAM and a dedicated GPU are recommended.

Costs stay modest with local execution since there are no API fees. You only pay for electricity and hardware. Cloud APIs can add up quickly with regular tests and larger test sets.

From a security perspective, local execution is an advantage. Your test questions, reference answers, and retrieved text passages stay on your system. This matters when your knowledge base contains sensitive company data.

Further Reading and Resources on Quality Testing

  • The official RAGAS documentation provides detailed explanations of all metrics and integration examples
  • The article on Reranking shows how to improve your retrieval precision
  • Hybrid Search combines two search methods for better results
  • When updating your knowledge base, the guide on Knowledge Base Updates will help
  • Fundamentals of local AI are covered in What is Local AI?

FAQ: Testing RAG Quality - Common Questions

How many questions should my test set contain? For an initial assessment, 20 questions suffice. For meaningful metrics, 50 or more are recommended. More questions give more reliable values, but also more work.

Do I need RAGAS or are manual tests enough? For getting started, manual tests work fine. They give you a feel for the system. RAGAS is worth it when you want systematic comparison, like before and after a change, or when managing larger systems.

Which model should I use as my RAGAS evaluator? A model with 7 to 8 billion parameters is a good starting point. Smaller models can be unreliable. If quality matters more than speed, use a larger model.

How often should I run tests? At minimum, after any significant change: knowledge base updates, parameter adjustments, or language model switches. With active development, testing after each sprint makes sense.

What if metrics look good but answers seem poor? That’s a sign of overfitting. Metrics measure specific aspects but not everything. Use manual spot checks and look at answers directly. Don’t trust the numbers blindly.

Can I use RAGAS without reference answers? RAGAS can calculate some metrics without reference answers, like Faithfulness and Answer Relevance. Context Recall requires them. Without them, the value of measurement is limited.

How do I tell retrieval errors apart from generation errors? Look at the retrieved text passages. If the right sources were found but the answer is wrong, the problem is in generation. If the wrong sources were found, it’s a retrieval issue. RAGAS separates this through Context Precision and Context Recall versus Faithfulness.

What’s a good Faithfulness score? Scores above 0.8 are solid, above 0.9 very good. Below 0.7 you should make urgent improvements, because the model is making things up too often. Keep in mind the score depends on your evaluation model and can fluctuate slightly between runs.

Do source citations help quality? Source citations make answers more transparent to users and help with manual review. They don’t directly improve metrics, but they make errors more visible.

What does a quality test with RAGAS on local hardware cost? Only electricity and time. With 50 questions on an 8B model on an average GPU, a run takes about 15 to 30 minutes. Without a GPU it can take much longer.

Can I automate tests? Yes, RAGAS integrates into Python scripts. You can run evaluation automatically after each knowledge base update and log the results. This lets you track trends over time.

Sources and Further Reading

  • RAGAS documentation and GitHub repository
  • Es, S., James, J., Espinosa-Anke, L., Schockaert, S. (2023): “RAGAS: Automated Evaluation of Retrieval Augmented Generation”
  • Lewis, P. et al. (2020): “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”
  • More articles on BotServ.de on local RAG
Back to Blog
Share:

Related Posts