Skip to main content
Built-in extractor names are a deprecated alias — collections are now created by picking features. This pipeline is selected with features: ["text_search"]. Existing feature_extractor configs keep working; see the migration guide.

View on GitHub

Runnable reference for this extractor — inputs, parameters, output fields, embedding models, and copy-paste examples. Auto-generated from the live registry.
Text extractor pipeline showing chunking, E5-Large embedding, and optional LLM extraction
The text extractor generates dense vector embeddings from text using the E5-Large multilingual model. Optimized for semantic search, RAG applications, and general-purpose text retrieval. Supports text chunking/decomposition with multiple splitting strategies. Fast (5ms/doc), cost-effective (free), and supports 100+ languages.
Cross-lingual retrieval. Because all 100+ languages share one embedding space, a query in one language matches content in any language — an English query (e.g. “foreign language”, “translate what they said”) retrieves Farsi, Mandarin, or Spanish text with no translation step and no per-language index. Note that this is retrieval, not translation: matched text is returned in its original language. To render it in the reader’s language, add an llm_enrich translate stage — see the Search Across Languages recipe.
View extractor details at api.mixpeek.com/v1/collections/features/extractors/text_extractor_v1 or fetch programmatically with GET /v1/collections/features/extractors/{feature_extractor_id}.

Pipeline Steps

  1. Filter Dataset (if collection_id provided)
    • Filter to specified collection
  2. Apply Input Mappings
    • Resolve text field from source (e.g., transcription, content, data)
  3. Text Chunking (conditional: if split_by != "none")
    • Split by: characters, words, sentences, paragraphs, or pages
    • Configure chunk_size and chunk_overlap
    • Each chunk becomes a separate document
  4. E5 Text Embedding Generation
    • Multilingual E5-Large model (1024D)
    • L2 normalized vectors
    • Batch size: 4,096 texts
  5. Output
    • Text documents with embeddings
    • One document per input (or per chunk if chunking enabled)

When to Use

When NOT to Use

Dense embeddings and exact-keyword matching are complementary. Keep text_extractor for semantic recall and add a lexical: true search over a text index for exact tokens — fuse them with rrf. See Lexical (BM25) Search.

Input Schema

Input Examples:

Output Schema

When chunking is enabled, each chunk becomes a separate document with tracking metadata stored in metadata (not in the document payload):
  • chunk_index – Position of this chunk in the original document
  • chunk_text – The text content of this chunk
  • total_chunks – Total number of chunks from the source

Parameters

Chunking Parameters

Split Strategies

Recommended chunk sizes:
  • characters: 500-2000
  • words: 100-400
  • sentences: 3-10
  • paragraphs: 1-3
  • pages: 1
Chunk overlap: 10-20% of chunk_size helps preserve context across boundaries. Example: chunk_size: 1000, chunk_overlap: 100-200.

Embedding Model

The text extractor resolves its model through the central embedding registry. Leave unset to use the current TEXT modality default (intfloat_e5_large_instruct_v1, 1024d) — the registry swaps hot when a new frontier text model ships, so existing collections pick it up without a code change.
Dimensions are locked at namespace creation. Switching embedding_model on an existing namespace requires a migration since the vector index dimensionality is fixed.

Embedding Task

Instruction-aware embedding models (E5, Gemini) use a task hint to optimize the embedding for a specific downstream use case. By default, all extractors use retrieval_document at ingestion time, which produces embeddings optimized for asymmetric search (queries find documents).
Set embedding_task at the collection level, not on the extractor. See Collection Embedding Task for full details and examples.
You almost never need to set this. The default retrieval_document is correct for search, and at query time Mixpeek automatically uses retrieval_query. Only override if your collection is primarily used for clustering, classification, or symmetric similarity — not retrieval.
Non-instruction-aware models (SigLIP, CLIP, Vertex multimodal) ignore this parameter.

LLM Structured Extraction Parameters

response_shape Modes

Natural Language Mode (string):
The service automatically infers JSON schema from your description. JSON Schema Mode (object):

LLM Provider & Model Options

With llm_provider: "typesafe", Jev answers each response_shape field as a typed question. response_shape must be a JSON schema, and every field needs a closed set of values: a boolean, an enum, a number from 0 to 1, or an integer with minimum and maximum. Natural-language response_shape needs a generative model. The collection is rejected at creation if a field cannot be answered, and the error names the field.

Configuration Examples

Performance & Costs

LLM extraction adds cost and latency based on provider pricing. Only use when structured extraction is needed.

Dense Embeddings vs Lexical (BM25)

text_extractor produces dense embeddings — great for semantic recall, weak on exact tokens. For exact-keyword precision, pair it with a lexical (BM25) search (the lexical: true option on a feature_search stage, backed by a text payload index — not a separate extractor).
The strongest setup is hybrid: run a dense search and a lexical search in the same feature_search stage and fuse with rrf. See Lexical (BM25) Search.

Vector Index

In retrievers, reference this feature by its Feature URI above (the output name is multilingual_e5_large_instruct_v1, not the index name text_extractor_v1_embedding).

Bring your own vectors

To store vectors you computed yourself in a text_extractor collection, embed the way the platform does. The platform L2-normalizes its own vectors. Normalize yours the same way so they compare with the managed ones. The model and dimensions above are the current text default. Check index.dimensions on the collection’s vector_indexes entry before you embed, because the default can change. Send each vector under a key the collection accepts. Write vectors into a managed collection lists both.

Limitations

  • Token limit: 512 tokens (~400 words). Longer text is automatically truncated.
  • Exact phrases: Cannot reliably match exact phrases or technical terms.
  • Domain jargon: Struggles with very domain-specific jargon or acronyms.
  • Terminology variance: May miss documents that use different terminology for the same concept.
  • Short texts: Less effective for very short texts (1-5 words) where lexical matching is sufficient.
  • Keyword-heavy queries: Less effective for queries like “iPhone 15 Pro Max 256GB”.