Advanced RAG beyond current templates
Advanced RAG adds techniques for measured failures.
What you will learn
- Explain: rag accesses evidence, it does not retrain
- Apply the idea in an example: Advanced RAG beyond current templates
- Recognize limitations and verify the exercise outcome
Advanced RAG adds techniques for measured failures. They are not automatically options in local templates.
This is a conceptual or external lab. It does not assume the application exposes every control described.
How it works, step by step
RAG accesses evidence, it does not retrain
RAG means retrieval-augmented generation: retrieval supplies relevant passages before the model answers. Ingestion prepares documents, while query execution searches the index and builds context. An embedding model represents text numerically; the response model writes the answer. Rerankers, vision, graph and router models have different roles when needed. RAG helps with private or changing information but does not guarantee truth. Use an authorized API for balances, exact inventory or permissions; an old document is not a reliable source for current state.
Search and ranking are different stages
Retrieval selects candidates and ranking orders them. top_k is the number of requested results; increasing it can bring evidence but also noise. A reranker examines the query and candidates more closely. It cannot recover a passage missing from the initial list. BM25 looks for lexical matches and helps with exact codes; vector search looks for semantic similarity. RRF combines positions in ranked lists instead of adding scores with incompatible units. Extra stages must justify their quality, latency and cost through testing.
Define success before optimizing
Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. These advanced retrieval controls are external concepts, not additional application templates. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
Query rewriting clarifies queries; multi-query searches variants; HyDE creates hypothetical retrieval text; contextual retrieval enriches chunks. Corrective/self-RAG checks and may repeat retrieval; agentic retrieval chooses steps dynamically.
Try it yourself
- Rewrite three ambiguous questions and check that intent remains unchanged.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
Retain the original query and test whether retrieval gains justify extra calls; hypothetical text is not evidence. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
Rewriting can change meaning or drop an exact identifier. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Does RAG change LLM parameters?
No. It supplies request-time context; parameter changes belong to training.
Can a reranker fix an unindexed document?
No. The passage must first exist in the index and candidate list.
What outcome should this exercise produce?
Retain the original query and test whether retrieval gains justify extra calls; hypothetical text is not evidence.
Words to remember
- Retrieval: Finding relevant information in accessible sources.
- Rerank: A more careful reordering of candidates already retrieved.
- Regresie / Regression: A change that breaks a previously correct case.
Sources and your next step
To prepare: D06 — Updates, deletion and reindexing