Skip to main content
POST
Export Collection

Authorizations

Authorization
string
header
required

Mixpeek API key, sent as Authorization: Bearer mxp_sk_.... Create one in Studio under Settings → API Keys, or with an admin key via POST /v1/organizations/users/{user_email}/api-keys. A missing header returns 403; an invalid or revoked key returns 401.

X-Namespace
string
header
required

Namespace id (ns_...), not the namespace name. This scopes the request rather than authenticating it, and it is required on every operation marked x-mixpeek-namespace-scoped.

Path Parameters

collection_identifier
string
required

The ID or name of the collection to export

Body

application/json

Request model for exporting collection data.

Export Formats:

  • JSON: Line-delimited JSON (JSONL) format, one document per line. Good for streaming and large files.
  • CSV: Comma-separated values. Best for tabular data analysis in spreadsheets.
  • PARQUET: Columnar format optimized for analytics. Best for large datasets and data pipelines.
  • WEBDATASET: Tar shards in the WebDataset layout, one sample per document. Each sample carries <key>.json (the same row the other formats emit) and, with include_media=True, the source object's bytes as <key>.<ext>. Read it with any WebDataset loader. samples_per_shard sets the shard size.

Vector Export: Vectors are stored separately from document metadata due to their large size. When include_vectors=True, vectors are exported to a separate file with the naming convention: {collection_name}_vectors.{format}

Lineage: When include_lineage=True, every row also carries lineage_chain (each processing step from the source object to this document), source_content_hash (SHA-256 of the source content) and document_created_at (ISO 8601). In JSON the chain is a list; in CSV and Parquet it is a JSON-encoded string so every row keeps one column type.

Field Selection: Use select_fields to export only specific fields, reducing file size for large collections. Supports dot notation for nested fields (e.g., "metadata.title").

Filtering: Apply filters to export a subset of documents. Uses the same LogicalOperator format as the documents list endpoint.

format
enum<string>
default:parquet

Export format: json (line-delimited), csv, or parquet (default).

Available options:
json,
csv,
parquet,
webdataset
include_vectors
boolean
default:false

Whether to include vectors in the export. Vectors are exported to a separate file due to their large size. This significantly increases export time and file size.

include_lineage
boolean
default:false

Add lineage columns to every row: lineage_chain (the processing steps from source object to document), source_content_hash (SHA-256 of the source content) and document_created_at (ISO 8601). Kept even when select_fields is set. In CSV and Parquet, lineage_chain is a JSON-encoded string.

select_fields
string[] | null

Specific fields to include in the export. If not provided, all fields are exported. Supports dot notation for nested fields (e.g., 'metadata.title', 'metadata.author').

Example:
filters
LogicalOperator · object | null

Filter conditions to export only matching documents. Uses LogicalOperator format (AND/OR/NOT) same as document listing.

sample_size
integer | null

Maximum number of documents to export. If not provided, exports all documents. Useful for testing exports or creating sample datasets.

Required range: 1 <= x <= 1000000
include_media
boolean
default:false

WebDataset only. Put the source object's bytes in each sample alongside its row, resolved through the document's root object. A document whose object is gone, or which never had one, ships with its row alone and is counted in the response's media summary. This moves the full media set, so expect a much larger export and a longer run.

samples_per_shard
integer
default:1000

WebDataset only. Documents per tar shard. Smaller shards parallelize better across training workers; larger shards mean fewer files to move.

Required range: 1 <= x <= 100000
destination
ExportDestination · object | null

Write the export into your own storage through a connection you own. The files and a manifest land under the prefix you give. The 7-day download copy is still written too, so the presigned URLs in the response keep working and nothing existing changes.

Response

Successful Response

Response model for collection export.

Contains the presigned URL for downloading the exported file. The URL is valid for a limited time (typically 1 hour).

download_url
string
required

Presigned URL for downloading the exported file. Valid for 1 hour.

s3_path
string
required

Full S3 path where the export is stored (for internal reference). For webdataset this is the prefix holding the shards, not a single file.

format
enum<string>
required

The format of the exported file.

Available options:
json,
csv,
parquet,
webdataset
document_count
integer
required

Number of documents included in the export.

Required range: x >= 0
file_size_bytes
integer
required

Size of the exported file in bytes. For webdataset this is every shard added together.

Required range: x >= 0
exported_at
string<date-time>
required

Timestamp when the export was completed.

vectors_download_url
string | null

Presigned URL for downloading the vectors file (if include_vectors=True). Vectors are exported separately due to their large size.

vectors_s3_path
string | null

Full S3 path for the vectors file (if include_vectors=True).

shard_count
integer | null

WebDataset only. Number of tar shards written. Always exact, even when the shards list below is capped.

Required range: x >= 0
shard_pattern
string | null

WebDataset only. Brace pattern naming every shard, e.g. documents-{000000..000007}.tar. Most loaders take this directly.

shards
ExportShard · object[] | null

WebDataset only. Per-shard paths and presigned URLs, capped at 200 entries. When shard_count exceeds that, address the rest through s3_path and shard_pattern.

media
ExportMediaSummary · object | null

WebDataset only, present when include_media was set. Counts what media reached the shards and what did not.

destination
ExportDestinationResult · object | null

Present when the request named a destination. Lists what was written into your storage and where the manifest is.