Twelve Labs builds a video understanding platform that lets enterprises search, analyse and extract insights from video content at scale. The company develops its own multimodal foundation models rather than assembling third-party components. Two flagship models anchor the technology: Marengo, a multimodal encoder foundation model, and Pegasus, a video-language foundation model. Technical work spans video understanding, computer vision, speech and audio understanding, natural language processing and the AI infrastructure required to serve these models.
The Video Understanding Platform is deployable on cloud, private cloud and on-premise environments, reflecting enterprise requirements around data residency. It processes petabytes of video data and is used by more than 30,000 developers and companies. Customers and use cases span media and entertainment, media production workflows, sports - including the NFL - enterprise software, and cloud and data infrastructure. The company holds a SOC 2 Type 2 certification.
Twelve Labs has raised $107m, with strategic investors including NVIDIA, NEA, Radical Ventures, Index Ventures, Snowflake and Databricks. Its work has been recognised by leading researchers for benchmark performance. The company is headquartered in San Francisco, with an APAC presence in Seoul, South Korea, and serves customers globally.






