Documentation

Ground every answer in your knowledge.

Select a retrieval architecture, configure its models and search limits, then index the documents that should ground responses.

6 retrieval patterns

Every RAG requires a name, embedding model, and vector dimension. Graph, router, vision, and synthesis models appear only when their capabilities are enabled.

Overview

Build a retrieval pipeline

A RAG configuration converts source documents into vectors and retrieves relevant context for each query. More advanced patterns add reranking, visual extraction, graph traversal, lexical search, or intent-based routing.

1. Architect
2. Configure
3. Ground

Step 01

Select a RAG pattern

Each pattern activates a different retrieval path and exposes only its relevant model and tuning controls in step two.

Naive RAG
Retrieve-and-Rerank
Multimodal RAG
Graph RAG
Hybrid RAG
Router RAG

Step 02 · Section 01

RAG Details

Identify the knowledge base and describe the domain covered by its indexed sources.

RAG Details section with name and description
RAG NameRequired

The display name used in RAG directories and selectors for agents, widgets, and API targets.

DescriptionOptional

A concise summary of the knowledge domain, source material, and retrieval purpose.

Step 02 · Section 02

Vector Embedding Model

The embedding model converts document chunks and queries into numeric vectors. The selected dimension fixes the vector shape used for indexing and similarity search.

Vector Embedding Model section with embedding and dimension controls
Embedding ModelRequired

Selects the provider model used to encode both document chunks and incoming queries.

Vector DimensionRequired

Sets the number of numeric values stored for each vector representation.

Step 02 · Section 03

Rerank AI Model

The rerank model applies a neural cross-encoder to score each candidate passage against the user query after broad vector retrieval. It re-orders the candidates so the final context keeps the most relevant passages.

Rerank AI Model section with cross-encoder selection and model details
Rerank ModelRequired

Selects the cross-encoder used to score vector-search candidates against each query.

Step 02 · Sections 04–07

Specialized AI Models

Additional models appear only when the architecture or an optional capability needs them. Each compatible reasoning model can also expose an effort selector after selection.

Knowledge Graph Extraction Model section
Query Intent Classifier Model section
Enabled Vision and Media Description Model section
Enabled Response Synthesis AI Model section
Graph ModelConditional

Extracts entity nodes, semantic relationships, and structured triples during indexing.

Router ModelConditional

Classifies incoming intent and selects vector retrieval, graph traversal, or direct response.

Vision ModelConditional

Describes images, charts, and diagrams so visual information can join the searchable context.

Response ModelConditional

Turns retrieved passages into a grounded conversational answer for standalone RAG queries and widgets.

Reasoning EffortConditional

Applies a supported effort level to graph, router, vision, or response models.

TemperatureConditional

Controls variation in generated output for applicable generative stages.

Step 02 · Section 08

Architecture & Retrieval Parameters

Sliders control the amount of candidate context collected at each stage. The card changes with the selected pattern, so only parameters used by that pipeline are shown.

Router RAG Architecture and Retrieval Parameters section
top_kNaive, Multimodal, Router5 · 1–50Passages injected into the final context.
initial_top_kRetrieve-and-Rerank25 · 5–100Broad semantic candidate pool before reranking.
final_top_kRerank, Hybrid5 · 1–20Passages retained after reranking or rank fusion.
vector_top_kGraph4 · 1–20Semantic seed chunks used to locate graph entities.
graph_depthGraph1 · 1–3Relationship hops traversed from the seed entities.
graph_limitGraph20 · 1–100Maximum entity-relation triples added to context.
vector_kHybrid15 · 1–50Dense semantic candidates included in rank fusion.
bm25_kHybrid15 · 1–50Sparse lexical candidates included in rank fusion.

Step 02 · Section 08

Knowledge Base Documents

Save the RAG first to enable document uploads. Uploaded files are indexed with the selected embedding model and become available to retrieval after processing completes.

Knowledge Base Documents section before the RAG is saved

Changing the embedding model or vector dimension after indexing can require documents to be reprocessed so stored vectors remain compatible.

Step 02 · Section 09

Live Query Testing

After saving and indexing documents, submit representative queries to inspect the grounded answer, retrieved chunks, extracted facts, execution steps, and assembled context.

Live Query Testing and Grounded Preview section before the RAG is saved

Required before save

Pattern, RAG name, embedding model, vector dimension, and every model required by the selected architecture.

After creation

Use questions that require exact facts and multi-source synthesis, then adjust retrieval limits based on the returned evidence.