View on GitHub
Runnable reference for this extractor — inputs, parameters, output fields, embedding models, and copy-paste examples. Auto-generated from the live registry.
View extractor details at api.mixpeek.com/v1/collections/features/extractors/universal_extractor_v1 or fetch programmatically with
GET /v1/collections/features/extractors/{feature_extractor_id}.Pipeline Steps
- Resolve input — apply
input_mappingsto get the file URL/path from the source object (contentfield). - Detect modality — classify the object as image, video, audio, or document.
- Segment (if needed) — video and audio are processed in up to
max_video_segments30s segments; documents up tomax_document_pagespages. - Gemini embedding — generate a 3072-d Gemini Embedding 2 vector (
output_dimensionalityconfigurable 256–3072). - Text extraction (if
extract_text) — OCR for images/documents, transcription for audio/video. - Description (if
generate_description) — Gemini vision/understanding produces a natural-language description. - Output — one document per object (or per segment/page for chunked content).
When to Use
When NOT to Use
Input Schema
Output Schema
Output by Modality
The extractor writes one document per unit: an image is one document, video and audio split into 30-second segments, and a document splits into pages. Every document carriesuniversal_extractor_v1_embedding, modality, text, description, segment_index, and segment_total. The table lists the fields each modality adds and what fills text and description.
covered_s is the number of seconds the indexed segments span, capped by duration_s. It shows when max_video_segments cut a long file short.
Two parameters change the fields. generate_description: false sets description to null. extract_text: false leaves OCR text, the transcript, and page text out of text.
If no segment of a video or audio file produces a document, the extractor writes one document for the whole file. That document has a null text and no start_time_s or end_time_s.
What the Extractor Does Not Emit
The output has no field for a hook segment, a music or tone descriptor, keyframes, scene boundaries, or faces. Video segments are fixed 30-second windows that start at 0. The first segment (segment_index 0) spans the first 30 seconds of the file and cannot be narrowed to the first 3. The video description is free text, so no music or tone value is searchable as its own field.
To get these fields, use one of two paths:
- Face data: the Face Identity Extractor detects and embeds faces.
- Hook, music, or tone fields: build a custom extractor that writes them. Upload and deploy run on dedicated infrastructure. On the shared API, submit the extractor through Extractor Submissions.
Parameters
Configuration Examples
Performance & Costs
Vector Index
Limitations
- External dependency: Requires Google Gemini API availability; subject to its rate limits.
- Per-object cost: Higher per-object cost than self-hosted single-modality extractors.
- Segment/page caps: Video beyond
max_video_segmentsand documents beyondmax_document_pagesare truncated. - Download ceiling: Files larger than
max_file_download_mbare skipped on the Celery fast-path.

