Why does RAG answer poorly?

Repair the failed stage.

What you will learn

  • Explain: good documents come before good answers
  • Apply the idea in an example: Why does RAG answer poorly?
  • Recognize limitations and verify the exercise outcome

Repair the failed stage. Prompt changes do not help when necessary evidence never reaches context.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Good documents come before good answers

Prepare readable text, clear headings, tables with headers and sources without unnecessary duplicates. Chunking divides a document into separately retrievable passages; overlap preserves continuity between neighbors. A short chunk can lose conditions, while a large one introduces noise. Metadata preserves source, position and version, plus access scope where enforced by the implementation. Upload stores the file; indexing makes it searchable. When a source changes, check that old chunks disappear and new ones are retrieved before declaring the knowledge base current.

Search and ranking are different stages

Retrieval selects candidates and ranking orders them. top_k is the number of requested results; increasing it can bring evidence but also noise. A reranker examines the query and candidates more closely. It cannot recover a passage missing from the initial list. BM25 looks for lexical matches and helps with exact codes; vector search looks for semantic similarity. RRF combines positions in ranked lists instead of adding scores with incompatible units. Extra stages must justify their quality, latency and cost through testing.

Fluency and truth are checked separately

A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.

The visual map

Why does RAG answer poorly? Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: Representative test questions → Semantic search; Semantic search → Actually retrieved evidence; Actually retrieved evidence → Compare against expected results; Compare against expected results → User revises the request; User revises the request → Playground tests; Source documents → Parse and clean text; Parse and clean text → Semantic search; Indexing status → Semantic search; Expected facts and sources → Compare against expected results; Final answer → Compare against expected results; User revises the request → Semantic search. D05 · Relationship map Why does RAG answer poorly? Input Representative test questions Input Source documents Data Indexing status Processing Parse and clean text Processing Semantic search Data Actually retrieved evidence Data Expected facts and sources Decision / control Compare against expected results Output Final answer Input User revises the request Processing Playground tests Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

Is the text in the PDF? Was it extracted? Does the chunk exist? Did it become a candidate? Is it in context? Does the answer use it correctly? These checks separate six causes.

Try it yourself

  1. Diagnose empty extraction, low ranking, truncated conditions, stale sources and invented claims.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

Identify the first stage with missing or incorrect data and test one targeted change. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

Increasing top_k cannot fix empty extraction. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does completed upload mean a document is ready for RAG?

No. Extraction, chunking and indexing must also complete successfully.

Can a reranker fix an unindexed document?

No. The passage must first exist in the index and candidate list.

What outcome should this exercise produce?

Identify the first stage with missing or incorrect data and test one targeted change.

Words to remember

  • Chunk: A document passage that can be retrieved separately.
  • Rerank: A more careful reordering of candidates already retrieved.
  • Grounding: Grounding an answer in verifiable evidence.

Sources and your next step

To prepare: D04 — Good questions for testing RAG