Rerank RAG: better evidence first

Rerank adds more careful selection after retrieval.

What you will learn

  • Explain: search and ranking are different stages
  • Apply the idea in an example: Rerank RAG: better evidence first
  • Recognize limitations and verify the exercise outcome

Rerank adds more careful selection after retrieval. It helps when good evidence exists but ranks too low.

These steps refer to features identified in the application code. This material has not been validated on a live instance; check the documentation and the options in your version.

How it works, step by step

Search and ranking are different stages

Retrieval selects candidates and ranking orders them. top_k is the number of requested results; increasing it can bring evidence but also noise. A reranker examines the query and candidates more closely. It cannot recover a passage missing from the initial list. BM25 looks for lexical matches and helps with exact codes; vector search looks for semantic similarity. RRF combines positions in ranked lists instead of adding scores with incompatible units. Extra stages must justify their quality, latency and cost through testing.

RAG accesses evidence, it does not retrain

RAG means retrieval-augmented generation: retrieval supplies relevant passages before the model answers. Ingestion prepares documents, while query execution searches the index and builds context. An embedding model represents text numerically; the response model writes the answer. Rerankers, vision, graph and router models have different roles when needed. RAG helps with private or changing information but does not guarantee truth. Use an authorized API for balances, exact inventory or permissions; an old document is not a reliable source for current state.

Cost belongs to the whole path

Count model calls, input and output tokens, embeddings, reranking, vision and tools. Reasoning can consume billable tokens without an equally long visible response. Indexing costs and conversation costs differ. Use current prices rather than a permanent number from a course. A simple estimate is tokens/1,000,000 × price plus additional operations; application credits may have their own conversion. Compare cost per successful task, not just per call. Parallel execution can reduce latency while increasing total consumption. A budget needs a verified stopping condition.

Detailed lab

Interpret initial and final lists

initial_top_k determines retrieved candidate count; final_top_k determines how many remain after reranking. Check the selected reranker and its cost. Delivery text may be semantically close to a return question without being appropriate evidence.

Compare runs using identical corpus, query and response model. If the relevant passage is absent from both initial lists, reranking cannot solve the cause. If present and better ranked, also inspect the final answer: a better list does not guarantee every condition is used.

The visual map

Rerank RAG: better evidence first Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. Connections: Source documents → Document ingestion and indexing; Document ingestion and indexing → Vector index; User request → Semantic search; Semantic search → Candidate passages; Candidate passages → Reranker: rescore relevance; Reranker: rescore relevance → Selected passages; Selected passages → Build evidence context; Build evidence context → Response model (optional); Response model (optional) → Evidence with sources; Vector index → Semantic search; User request → Reranker: rescore relevance; Instructions and context → Response model (optional). R04 · Relationship map Rerank RAG: better evidence first Input Source documents Processing Document ingestion and indexing Store / index Vector index Input User request Processing Semantic search Data Candidate passages AI model / agent Reranker: rescore relevance Data Selected passages Processing Build evidence context Data Instructions and context AI model / agent Response model (optional) Output Evidence with sources Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

The return policy is the eighth candidate. initial_top_k=15 includes it; final_top_k=5 keeps the most relevant reranked passages. With only four initial candidates, it would be absent.

template: Rerank
initial_top_k: 15
final_top_k: 5

Try it yourself

  1. Configure a reranker, initial_top_k=15 and final_top_k=5 as an experiment, not universal settings.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

Correct evidence is included and better ordered; compare quality and cost with Naive. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

Reranking cannot fix unindexed or unauthorized sources. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Can a reranker fix an unindexed document?

No. The passage must first exist in the index and candidate list.

Does RAG change LLM parameters?

No. It supplies request-time context; parameter changes belong to training.

What outcome should this exercise produce?

Correct evidence is included and better ordered; compare quality and cost with Naive.

Words to remember

  • Rerank: A more careful reordering of candidates already retrieved.
  • Retrieval: Finding relevant information in accessible sources.
  • Latență / Latency: Time until a useful result.

Sources and your next step

To prepare: R03 — Naive RAG: your baseline