Hybrid RAG: meaning and exact words
Hybrid combines semantic and lexical search.
What you will learn
- Explain: search and ranking are different stages
- Apply the idea in an example: Hybrid RAG: meaning and exact words
- Recognize limitations and verify the exercise outcome
Hybrid combines semantic and lexical search. It helps queries mixing exact identifiers and natural-language descriptions.
These steps refer to features identified in the application code. This material has not been validated on a live instance; check the documentation and the options in your version.
How it works, step by step
Search and ranking are different stages
Retrieval selects candidates and ranking orders them. top_k is the number of requested results; increasing it can bring evidence but also noise. A reranker examines the query and candidates more closely. It cannot recover a passage missing from the initial list. BM25 looks for lexical matches and helps with exact codes; vector search looks for semantic similarity. RRF combines positions in ranked lists instead of adding scores with incompatible units. Extra stages must justify their quality, latency and cost through testing.
Coordinates for meaning
An embedding is a numeric vector computed by a model. Texts with related meanings tend to lie near each other in that model’s space, but similarity does not guarantee relevance. Queries and documents must share a compatible space: matching model and dimensions. A two-axis map is illustrative; real vectors can have hundreds or thousands of dimensions. Rebuild and retest the index when changing model or dimensions. Cosine similarity compares vector directions; it is not the probability that an answer is true.
Define success before optimizing
Prepare representative requests and expected answers or properties. Include ordinary cases, ambiguity, missing data and tool errors. For RAG, measure evidence retrieval separately from answer correctness. Recall@k measures the share of relevant evidence found in the first k results; precision@k measures the share of returned results that are relevant. For agents, also track actions, permissions and stopping. An LLM-as-judge can speed evaluation but needs calibration against humans. Do not change prompt, model and index simultaneously if you want to understand what improved results.
Detailed lab
What does RRF combine?
RRF favors documents ranked highly in one or more lists. It uses rank rather than assuming BM25 scores and vector similarity share a scale. In the local template the fused list is reranked before context assembly.
For example, vector search places the manual above the catalog; BM25 places the LX-240 catalog first. Fusion can retain both evidence sources. Reranking checks query relevance. If neither list contains required evidence, fusion cannot create it.
The visual map
Follow the solid arrows for the main flow. Dashed blue arrows supply data or context; dashed pink arrows show feedback or returning results. Colors and shapes distinguish models, stores, decisions and outputs. Semantic and BM25 candidate lists are combined by RRF, then reranked. The two searches are complementary; the drawing does not imply concurrent execution. Retrieval returns evidence. If a response model is enabled, it also composes an answer; retrieved evidence still needs checking. On smaller screens, scroll horizontally to follow the entire diagram.
A complete example
For “adapter for LX-240”, vector search captures meaning and BM25 helps with the exact identifier. Local Hybrid applies RRF, then rerank and generation.
template: Hybrid
vector_k: 15
bm25_k: 15
final_top_k: 5
Try it yourself
- Configure vector_k=15, bm25_k=15 and final_top_k=5; compare with Naive on exact identifiers.
- Record the input, source and expected outcome before running the experiment. Use only the fictional data in the example.
- Follow the diagram stages. At every step record what information is received and produced; do not confuse intermediate output with the final outcome.
- Repeat after removing necessary information or making the input ambiguous. Check whether the system clarifies, stops or invents an answer.
- Compare with the explained solution. Keep the configuration, date, result and an explanation for differences. Change one thing and retest.
An explained solution
The right passage reaches the final list without directly adding incompatible vector and BM25 scores. A successful exercise lets you show the connection between input, stages and outcome. When information is missing, a cautious answer is more useful than invented details. Compare more than style: check conditions, sources and operations too.
When it helps and what can go wrong
An exact lexical match may still come from the wrong version. Choose this approach when it improves a measured need. Keep a simple baseline and compare outcomes using identical inputs. One successful example does not establish reliability in every situation.
Check your understanding
Can a reranker fix an unindexed document?
No. The passage must first exist in the index and candidate list.
Can you directly compare vectors from different models?
Generally no; their spaces are not automatically compatible.
What outcome should this exercise produce?
The right passage reaches the final list without directly adding incompatible vector and BM25 scores.
Words to remember
- Rerank: A more careful reordering of candidates already retrieved.
- Embedding: A numeric representation used to compare content.
- Regresie / Regression: A change that breaks a previously correct case.
Sources and your next step
To prepare: R06 — Graph RAG and knowledge graphs