Embeddings and vector search
Embeddings enable search by meaning even when the question differs from document wording.
What you will learn
- Explain: coordinates for meaning
- Apply the idea in an example: Embeddings and vector search
- Recognize limitations and verify the exercise outcome
Embeddings enable search by meaning even when the question differs from document wording. Vectors are model-specific coordinates.
You can follow this course without an account or programming. Examples use synthetic data and AI outcomes need checking.
How it works, step by step
Coordinates for meaning
An embedding is a numeric vector computed by a model. Texts with related meanings tend to lie near each other in that model’s space, but similarity does not guarantee relevance. Queries and documents must share a compatible space: matching model and dimensions. A two-axis map is illustrative; real vectors can have hundreds or thousands of dimensions. Rebuild and retest the index when changing model or dimensions. Cosine similarity compares vector directions; it is not the probability that an answer is true.
Search and ranking are different stages
Retrieval selects candidates and ranking orders them. top_k is the number of requested results; increasing it can bring evidence but also noise. A reranker examines the query and candidates more closely. It cannot recover a passage missing from the initial list. BM25 looks for lexical matches and helps with exact codes; vector search looks for semantic similarity. RRF combines positions in ranked lists instead of adding scores with incompatible units. Extra stages must justify their quality, latency and cost through testing.
Good documents come before good answers
Prepare readable text, clear headings, tables with headers and sources without unnecessary duplicates. Chunking divides a document into separately retrievable passages; overlap preserves continuity between neighbors. A short chunk can lose conditions, while a large one introduces noise. Metadata preserves source, position and version, plus access scope where enforced by the implementation. Upload stores the file; indexing makes it searchable. When a source changes, check that old chunks disappear and new ones are retrieved before declaring the knowledge base current.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
“When does my parcel arrive?” can find “Delivery takes 2–3 days”. “Code LX-240” also has an exact-match component that lexical search can help.
Try it yourself
- Compare two equivalent questions and one with a different exact code.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
Identify the relevant passage and check its conditions; similarity is not truth probability. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
Equal dimensions do not guarantee compatible vector spaces. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Can you directly compare vectors from different models?
Generally no; their spaces are not automatically compatible.
Can a reranker fix an unindexed document?
No. The passage must first exist in the index and candidate list.
What outcome should this exercise produce?
Identify the relevant passage and check its conditions; similarity is not truth probability.
Words to remember
- Embedding: A numeric representation used to compare content.
- Rerank: A more careful reordering of candidates already retrieved.
- Chunk: A document passage that can be retrieved separately.
Sources and your next step
To prepare: R01 — What is RAG and why is it useful?