Chunking and metadata
Chunking determines how much context is retrieved in one passage.
What you will learn
- Explain: good documents come before good answers
- Apply the idea in an example: Chunking and metadata
- Recognize limitations and verify the exercise outcome
Chunking determines how much context is retrieved in one passage. Metadata tells us where it came from and whether it is current.
This is a conceptual or external lab. It does not assume the application exposes every control described.
How it works, step by step
Good documents come before good answers
Prepare readable text, clear headings, tables with headers and sources without unnecessary duplicates. Chunking divides a document into separately retrievable passages; overlap preserves continuity between neighbors. A short chunk can lose conditions, while a large one introduces noise. Metadata preserves source, position and version, plus access scope where enforced by the implementation. Upload stores the file; indexing makes it searchable. When a source changes, check that old chunks disappear and new ones are retrieved before declaring the knowledge base current.
Coordinates for meaning
An embedding is a numeric vector computed by a model. Texts with related meanings tend to lie near each other in that model’s space, but similarity does not guarantee relevance. Queries and documents must share a compatible space: matching model and dimensions. A two-axis map is illustrative; real vectors can have hundreds or thousands of dimensions. Rebuild and retest the index when changing model or dimensions. Cosine similarity compares vector directions; it is not the probability that an answer is true.
Define success before optimizing
Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
Separating “14 days” from “unused only” loses the condition. Keep rules and exceptions together; retain table headers in extracted representations.
Try it yourself
- Split the policy two ways and compare information lost at boundaries.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
The relevant passage preserves deadline and condition, plus source identifier. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
Do not assume a chunk-size UI control exists unless exposed. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Does completed upload mean a document is ready for RAG?
No. Extraction, chunking and indexing must also complete successfully.
Can you directly compare vectors from different models?
Generally no; their spaces are not automatically compatible.
What outcome should this exercise produce?
The relevant passage preserves deadline and condition, plus source identifier.
Words to remember
- Chunk: A document passage that can be retrieved separately.
- Embedding: A numeric representation used to compare content.
- Regresie / Regression: A change that breaks a previously correct case.
Sources and your next step
To prepare: D02 — Upload, fast indexing and batch indexing