Skip to main content
Cluster stage showing semantic grouping of search results
The Cluster stage groups documents based on embedding similarity, creating semantic clusters of related content. This helps organize search results into meaningful groups and discover themes within your results.
Stage Category: GROUP (Groups documents)Transformation: N documents → K clusters with documents

When to Use

When NOT to Use

Parameters

Clustering Algorithms

Documents that hdbscan or dbscan treat as noise get the cluster ID -1. The clusters and representatives output modes leave them out, and labeled keeps them with cluster_id: -1.

Configuration Examples

How Clustering Works

  1. Extract Embeddings: Get embedding vectors from each document
  2. Apply Algorithm: Run the clustering algorithm (e.g., k-means)
  3. Assign Documents: Each document gets a cluster ID
  4. Compute Centroids: Calculate cluster centers
  5. Format Output: Return cluster documents, labeled documents, or one representative per cluster, depending on output_mode
The stage needs at least 3 documents with an embedding. With fewer, it returns no results and reports the reason under skipped in the stage metadata.

Output Schema

The shape depends on output_mode.

clusters (default)

The stage returns one document per cluster. Each has a generated document_id, the collection_id of the source documents, and a score equal to the member count. The members and the centroid are carried in payload.
representative_id is the member closest to the centroid. centroid holds the vector only when include_centroids is true and the query requests vectors. Otherwise it is null. At most max_members_per_cluster members are listed, and member_count counts all of them.

labeled

The stage returns the original documents with two fields added.
cluster_distance is the distance to the cluster centroid, so lower is closer. It is absent when the algorithm returns no centroid, and for noise documents (cluster_id: -1).

representatives

The stage returns one document per cluster, the one closest to the centroid.
The stage metadata reports algorithm, n_clusters, n_documents_in, n_documents_out, n_valid_embeddings and n_invalid_documents.

Performance

Clustering large document sets (over 1,000) can be slow. Consider pre-filtering or sampling before clustering.

Common Pipeline Patterns

Search + Cluster

Cluster + Sample Representatives

Theme Discovery Pipeline

Diverse Results Pipeline

Choosing n_clusters

Start with fewer clusters and increase if clusters are too broad. Use hdbscan if you don’t know the optimal number.

Error Handling