Skip to main content
Video Datalake
One embedding model

Collections: one embedding model

Create Collection now takes an optional embedding_model. It defaults to omni_retriever — OmniRetriever, the 3072-dim multilingual vector space — and is currently the only accepted value, so you can keep omitting the field entirely.
  • Per-collection model selection is retired: every new collection indexes into the same vector space, and live streams no longer need a collection created a particular way.
  • Collections created before this change keep the model they were created with, so Move Video across two different models still returns 409.
Video Datalake
Video Datalake — preview

Video Datalake

First preview of the video data lake — 41 endpoints:
  • Collections with fixed embedding (OmniRetriever), optional safety detection and face recognition.
  • Ingestion — upload by URL, file, or resumable session; attach live streams.
  • Async operations — uniform 202 + Operation model with polling and HMAC-signed webhooks.
  • Moments & derived content — captions, transcripts, frames, clips, titles, summaries, speakers, entities.
  • Search — semantic, keyword, hybrid, and image, returning moment lists.
  • Face library and safety-detection events.
Embeddings standardized on OmniRetriever (3072-dim, multilingual); captions & summaries by OmniCaptioner.
MCPCLIPlugin
Agents & tools

Agents & tools

Operate the Datalake without writing glue code: