Skip to main content
Agent Search stage showing LLM reasoning loop with retriever stages as tools
The Agent Search stage uses an LLM reasoning loop to orchestrate other retriever stages as callable tools. Instead of executing a fixed sequence of stages, the LLM dynamically decides which stages to invoke, with what arguments, and how many iterations to perform based on the query and intermediate results.
Stage Category: FILTER (Adaptive retrieval)Transformation: Query → LLM reasoning (1-N iterations) → Refined documents

When to Use

When NOT to Use

Parameters

Strategies

Cluster Navigation

cluster_navigation searches a cluster hierarchy the way a person browses one. A cluster run writes one centroid document per cluster, with its label, summary and parent_cluster_id. The stage reads those centroids, starts at the top-level clusters, and asks the model which child cluster most likely contains documents that answer the query. It repeats this until it reaches leaf clusters. It then returns the member documents of the best leaf clusters. The walk is a beam search, and every decision goes to a decision model. Each decision returns a probability for every child cluster. The stage keeps the best beam_width paths and ranks them by the geometric mean of their decision probabilities, so shallow and deep leaves compare fairly. All paths at one level are asked in one request. model_name picks the decision model. Leave it unset to use the deployment default, which is jev-latest. Jev returns calibrated probabilities. A generative model such as gemini-2.5-flash-lite reports its own confidence for each option, and the stage marks those as uncalibrated in probabilities_calibrated. A parent cluster is shown to the model as the leaves a descent through it can reach, such as Sports: Skateboarding / Surfing. Labeling picks one label per cluster, so a parent that holds species and landforms can come back labeled Species. Listing the leaves keeps a query about a mountain from skipping that branch. The stage matches members to clusters by membership_field, so cluster ids must be unique within the collections it reads. Point centroid_collection_ids at one cluster’s output collection. If several cluster runs write cl_0 into the same collection, members of every cl_0 come back. When the stage receives documents from an earlier stage, it keeps the ones in the chosen clusters and preserves their order. When it runs first, it reads the members from the collections. Each returned document carries cluster_path and cluster_path_score. Stage metadata includes chosen_clusters (each leaf with its path and score), navigation_trace (the candidates and probabilities at every level), model_used and probabilities_calibrated.
See Hierarchical search over clusters for an end-to-end example.

Available Tools (Stages)

The agent can invoke any registered retriever stage as a tool. Each stage is presented to the LLM with a simplified parameter schema:

Configuration Examples

full_catalog hands the LLM every registered filter/sort/rerank/enrich stage as a tool and lets it decide the pipeline shape per query. Higher cost and latency than the pinned strategies — use it when composition quality matters more than cost.
Use the custom strategy when you need precise control over which tools the agent can access and how it should reason. The built-in strategies provide sensible defaults for common patterns.
Use feedback to create a human-in-the-loop refinement cycle: execute a retriever, review the results, then re-execute with feedback describing what was wrong. The agent adjusts its strategy based on your corrections. Pair with min_confidence to ensure the agent keeps searching until results meet your quality bar.

How It Works

  1. Strategy selection: If auto_strategy is enabled, a lightweight LLM call picks the best strategy for the query
  2. Prompt assembly: The system prompt is built from the strategy defaults, with feedback prepended if provided
  3. Reasoning loop: Each iteration, the LLM receives the query, available tools, and a budget note showing remaining iterations and seconds
  4. Tool execution: The LLM calls retriever stages as tools. Results are summarized and fed back
  5. Confidence gating: The LLM calls finish_search to declare done with a confidence score (0.0–1.0). If confidence is below min_confidence, the loop continues
  6. Context compression: When the conversation grows long, older messages are replaced with a compact working-memory summary to stay within context limits
  7. Return: Accumulated results and metadata (confidence, summary, reasoning trace) are returned
Each tool call creates a sub-state execution of the target stage, inheriting namespace and collection context from the parent. Non-empty results from each iteration replace the previous working set. If a refinement query returns zero results, the previous results are preserved.

Performance

Common Pipeline Patterns

Agent as First Stage

Complex Query Decomposition

Response Metadata

The stage returns execution metadata in stage_statistics:

Error Handling