Preparing documents for AI

A good knowledge base starts with clean, current documents.

What you will learn

  • Explain: good documents come before good answers
  • Apply the idea in an example: Preparing documents for AI
  • Recognize limitations and verify the exercise outcome

A good knowledge base starts with clean, current documents. The model cannot independently decide which conflicting policy is valid.

You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.

How it works, step by step

Good documents come before good answers

Prepare readable text, clear headings, tables with headers and sources without unnecessary duplicates. Chunking divides a document into separately retrievable passages; overlap preserves continuity between neighbors. A short chunk can lose conditions, while a large one introduces noise. Metadata preserves source, position and version, plus access scope where enforced by the implementation. Upload stores the file; indexing makes it searchable. When a source changes, check that old chunks disappear and new ones are retrieved before declaring the knowledge base current.

Fluency and truth are checked separately

A hallucination is unsupported or incorrect content presented as an answer. It can arise from missing information, ambiguity or incorrectly combined patterns. Ask for evidence and check the original document, date, units and conditions. A real citation may still fail to support the claim. Use current sources for current facts, a calculator for arithmetic, and “cannot determine” when evidence is missing. Do not treat confidence expressed in prose as a calibrated probability.

Check access before use

Permissions must be enforced by servers and the data layer. Do not let the model decide whether a user may see a document. Filter sources before they enter context, not after generating the answer. Use minimal-access credentials, keep them out of prompts, screenshots and browser code, and rotate exposed keys. A widget domain allowlist does not replace API authentication. Authentication identifies you; authorization determines the operations and data you can use. Test denied access too.

The visual map

Preparing documents for AI Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Connections: Source documents → Parse and clean text; Parse and clean text → Extracted text; Extracted text → Passages and source metadata; Passages and source metadata → Document ingestion and indexing; Document ingestion and indexing → Human verification; Consent and usage rights → Parse and clean text; Source, page and version → Passages and source metadata; Original sources → Human verification; Images, charts and tables → Parse and clean text. D01 · Relationship map Preparing documents for AI Input Source documents Decision / control Consent and usage rights Data Source, page and version Processing Parse and clean text Data Images, charts and tables Data Extracted text Data Passages and source metadata Store / index Original sources Processing Document ingestion and indexing Decision / control Human verification Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

An old PDF says 30 days; current policy says 14. Mark versions, retain the current source and remove the expired duplicate from the active corpus.

Try it yourself

  1. Clean two conflicting policies and a table without headers.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

The active corpus contains an identifiable policy without conflicting versions. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

Do not upload sensitive documents merely for an experiment. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does completed upload mean a document is ready for RAG?

No. Extraction, chunking and indexing must also complete successfully.

Is a link in an answer enough?

No; it must support the specific claim and fit the time and circumstances.

What outcome should this exercise produce?

The active corpus contains an identifiable policy without conflicting versions.

Words to remember

  • Chunk: A document passage that can be retrieved separately.
  • Grounding: Grounding an answer in verifiable evidence.
  • Autorizare / Authorization: Checking the right to access an operation or resource.

Sources and your next step

To prepare: R01 — What is RAG and why is it useful? · R03 — Naive RAG: your baseline