Why can AI be confidently wrong?

A well-written answer can hide a wrong fact.

What you will learn

  • Explain: fluency and truth are checked separately
  • Apply the idea in an example: Why can AI be confidently wrong?
  • Recognize limitations and verify the exercise outcome

A well-written answer can hide a wrong fact. Verification starts with asking what evidence supports each claim.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Fluency and truth are checked separately

A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.

RAG accesses evidence, it does not retrain

RAG means retrieval-augmented generation: retrieval supplies relevant passages before the model answers. Ingestion prepares documents, while query execution searches the index and builds context. An embedding model represents text numerically; the response model writes the answer. Rerankers, vision, graph and router models have different roles when needed. RAG helps with private or changing information but does not guarantee truth. Use an authorized API for balances, exact inventory or permissions; an old document is not a reliable source for current state.

Define success before optimizing

Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.

The visual map

Why can AI be confidently wrong? Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: User request → Agent / LLM; Agent / LLM → Claims in the answer; Claims in the answer → Human verification; Human verification → Result to verify; Original sources → Human verification; Calculator or executable check → Human verification; Missing or ambiguous information → Clarify or state uncertainty (Insufficient information); Clarify or state uncertainty → Agent / LLM; Human verification → Agent / LLM (Correct unsupported claims). F05 · Relationship map Why can AI be confidently wrong? Input User request Decision / control Missing or ambiguous information Store / index Original sources AI model / agent Agent / LLM Data Claims in the answer Decision / control Human verification Output Clarify or state uncertainty Tool / service Calculator or executable check Output Result to verify Insufficient information Correct unsupported claims Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

The fictional policy says 14 days for an unused product. The answer says “30 days, no conditions”. Check both the number and omitted condition against a valid source.

Try it yourself

  1. Label each claim as supported, contradicted or unsupported.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

A correct answer preserves 14 days and the unused condition; other situations need clarification. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

RAG reduces missing context but does not eliminate all hallucinations. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Is a link in an answer enough?

No; it must support the specific claim and fit the time and circumstances.

Does RAG change LLM parameters?

No. It supplies request-time context; parameter changes belong to training.

What outcome should this exercise produce?

A correct answer preserves 14 days and the unused condition; other situations need clarification.

Words to remember

  • Grounding: Grounding an answer in verifiable evidence.
  • Retrieval: Finding relevant information in accessible sources.
  • Regresie / Regression: A change that breaks a previously correct case.

Sources and your next step

To prepare: F04 — Context and memory: what does AI see?