Naive RAG: your baseline

Naive RAG is a clear baseline: semantic retrieval, context assembly and generation.

What you will learn

  • Explain: rag accesses evidence, it does not retrain
  • Apply the idea in an example: Naive RAG: your baseline
  • Recognize limitations and verify the exercise outcome

Naive RAG is a clear baseline: semantic retrieval, context assembly and generation. Low complexity makes diagnosis easier.

These steps refer to features identified in the application code. This material has not been validated on a live instance; check the documentation and the options in your version.

How it works, step by step

RAG accesses evidence, it does not retrain

RAG means retrieval-augmented generation: retrieval supplies relevant passages before the model answers. Ingestion prepares documents, while query execution searches the index and builds context. An embedding model represents text numerically; the response model writes the answer. Rerankers, vision, graph and router models have different roles when needed. RAG helps with private or changing information but does not guarantee truth. Use an authorized API for balances, exact inventory or permissions; an old document is not a reliable source for current state.

Coordinates for meaning

An embedding is a numeric vector computed by a model. Texts with related meanings tend to lie near each other in that model’s space, but similarity does not guarantee relevance. Queries and documents must share a compatible space: matching model and dimensions. A two-axis map is illustrative; real vectors can have hundreds or thousands of dimensions. Rebuild and retest the index when changing model or dimensions. Cosine similarity compares vector directions; it is not the probability that an answer is true.

Define success before optimizing

Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.

Detailed lab

Test retrieval before writing

Use the same catalog, delivery and policy fixtures. After indexing, ask “What is the deadline for an unused product?”. Retrieval should include the deadline and conditions. Then check the answer preserves them without inventing exceptions.

Ask “What commercial warranty applies?”. The corpus contains no such information. A good outcome acknowledges that absence. If retrieval misses policy, diagnose source and index; if it finds policy but the answer ignores it, inspect context and generation.

The visual map

Naive RAG: your baseline Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. Connections: Source documents → Passages and source metadata; Passages and source metadata → Embedding model; Embedding model → Vector index; User request → Query embedding; Query embedding → Semantic search; Semantic search → Candidate passages; Candidate passages → Build evidence context; Build evidence context → Response model (optional); Response model (optional) → Evidence with sources; Vector index → Semantic search; User request → Build evidence context; Instructions and context → Response model (optional). R03 · Relationship map Naive RAG: your baseline Input Source documents Data Passages and source metadata AI model / agent Embedding model Store / index Vector index Input User request AI model / agent Query embedding Processing Semantic search Data Candidate passages Data Instructions and context Processing Build evidence context AI model / agent Response model (optional) Output Evidence with sources Main flow Data and context Feedback and return

Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. On smaller screens, scroll horizontally to follow the entire diagram.

A complete example

You have policy, catalog and delivery documents. Start with top_k=4 and ask the return deadline. Check that the correct passage reaches context before changing the prompt.

template: Naive
top_k: 4

Try it yourself

  1. Create Naive RAG, choose embedding and response models, upload data and test after indexing.
  2. Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
  3. Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
  4. Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
  5. Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.

An explained solution

The policy passage is retrieved and the answer preserves 14 days and conditions. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.

When it helps and what can go wrong

Larger top_k can also add irrelevant context. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.

Check your understanding

Does RAG change LLM parameters?

No. It supplies request-time context; parameter changes belong to training.

Can you directly compare vectors from different models?

Generally no; their spaces are not automatically compatible.

What outcome should this exercise produce?

The policy passage is retrieved and the answer preserves 14 days and conditions.

Words to remember

  • Retrieval: Finding relevant information in accessible sources.
  • Embedding: A numeric representation used to compare content.
  • Regresie / Regression: A change that breaks a previously correct case.

Sources and your next step

To prepare: R02 — Embeddings and vector search