What is RAG and why is it useful?

RAG is like a colleague opening a folder before answering.

What you will learn

  • Explain: rag accesses evidence, it does not retrain
  • Apply the idea in an example: What is RAG and why is it useful?
  • Recognize limitations and verify the exercise outcome

RAG is like a colleague opening a folder before answering. Check that it finds the right page and preserves its meaning.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

RAG accesses evidence, it does not retrain

RAG means retrieval-augmented generation: retrieval supplies relevant passages before the model answers. Ingestion prepares documents, while query execution searches the index and builds context. An embedding model represents text numerically; the response model writes the answer. Rerankers, vision, graph and router models have different roles when needed. RAG helps with private or changing information but does not guarantee truth. Use an authorized API for balances, exact inventory or permissions; an old document is not a reliable source for current state.

Good documents come before good answers

Prepare readable text, clear headings, tables with headers and sources without unnecessary duplicates. Chunking divides a document into separately retrievable passages; overlap preserves continuity between neighbors. A short chunk can lose conditions, while a large one introduces noise. Metadata preserves source, position and version, plus access scope where enforced by the implementation. Upload stores the file; indexing makes it searchable. When a source changes, check that old chunks disappear and new ones are retrieved before declaring the knowledge base current.

Fluency and truth are checked separately

A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.

The visual map

What is RAG and why is it useful? Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. Connections: Source documents → Document ingestion and indexing; Document ingestion and indexing → Vector index; User request → Semantic search; Semantic search → Candidate passages; Candidate passages → Build evidence context; Build evidence context → Response model (optional); Response model (optional) → Evidence with sources; Evidence with sources → Human verification; Vector index → Semantic search; Instructions and context → Response model (optional). R01 · Relationship map What is RAG and why is it useful? Input Source documents Processing Document ingestion and indexing Store / index Vector index Input User request Processing Semantic search Data Candidate passages Data Instructions and context Processing Build evidence context AI model / agent Response model (optional) Output Evidence with sources Decision / control Human verification Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

Upload and index the fictional policy, then ask its deadline. Retrieval supplies the 14-day passage to the model. A warranty question absent from the source must be treated as missing information.

Try it yourself

  1. Draw ingestion and query execution separately for the return policy.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

The answer can be checked against the passage; unsupported questions receive no invented promise. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

Indexing and answering are separate paths with different failure modes. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does RAG change LLM parameters?

No. It supplies request-time context; parameter changes belong to training.

Does completed upload mean a document is ready for RAG?

No. Extraction, chunking and indexing must also complete successfully.

What outcome should this exercise produce?

The answer can be checked against the passage; unsupported questions receive no invented promise.

Words to remember

  • Retrieval: Finding relevant information in accessible sources.
  • Chunk: A document passage that can be retrieved separately.
  • Grounding: Grounding an answer in verifiable evidence.

Sources and your next step

To prepare: F03 — What is an LLM and how does it answer? · P02 — Context engineering: what reaches the model? · P07 — Prompt, RAG or fine-tuning?