Evaluating RAG stage by stage
Separate retrieved evidence from how the model uses it.
What you will learn
- Explain: define success before optimizing
- Apply the idea in an example: Evaluating RAG stage by stage
- Recognize limitations and verify the exercise outcome
Separate retrieved evidence from how the model uses it. Otherwise you may optimize the wrong stage.
You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.
How it works, step by step
Define success before optimizing
Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.
Search and ranking are different stages
Retrieval selects candidates and ranking orders them. top_k is the number of requested results; increasing it can bring evidence but also noise. A reranker examines the query and candidates more closely. It cannot recover a passage missing from the initial list. BM25 looks for lexical matches and helps with exact codes; vector search looks for semantic similarity. RRF combines positions in ranked lists instead of adding scores with incompatible units. Extra stages must justify their quality, latency and cost through testing.
Fluency and truth are checked separately
A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
Retrieve one of two relevant passages in four results: recall=1/2 and precision=1/4 for that case. The answer may still omit a condition from the retrieved passage.
Try it yourself
- Label relevant passages for five questions and compare retrieval with generation.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
Report evidence coverage, result relevance and final-claim support separately. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
A high vector score does not replace relevance labeling. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Does a good style score prove agent success?
No. The agent may write beautifully after using the wrong tool or data.
Can a reranker fix an unindexed document?
No. The passage must first exist in the index and candidate list.
What outcome should this exercise produce?
Report evidence coverage, result relevance and final-claim support separately.
Words to remember
- Regresie / Regression: A change that breaks a previously correct case.
- Rerank: A more careful reordering of candidates already retrieved.
- Grounding: Grounding an answer in verifiable evidence.
Sources and your next step
- LangGraph — Workflows and agents
- Lewis et al. — Retrieval-Augmented Generation
- Hugging Face — LLM Course
To prepare: E01 — How to measure whether AI is useful