Twelve Labs video AI just pulled in $100 million because investors are making AI startup bets that the next AI interface won’t be a chat box, it will be searchable footage. The startup said Wednesday, July 1, that it raised a $100 million Series B to expand its work on models that can understand, index, retrieve, and reason over video, according to PYMNTS.

Twelve Labs Grabs $100M as Video AI Battles Chatbots
XOOMAR Intelligence
Analyst Take
The round was co-led by NEA and NAVER Ventures, with participation from Amazon, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital, and Red Bull Ventures, according to TwelveLabs’ own announcement. The company did not disclose a valuation in the supplied materials.
“Five years ago, we began with a simple observation: The world does not happen in text. It happens in motion,” Co-Founder and CEO Jae Lee wrote.
Twelve Labs lands $100 million to scale its video AI platform
The core thesis behind Twelve Labs video AI is blunt: text has been made programmable, but video remains largely locked away from machines. Lee argues that most AI systems still work from compressed descriptions of reality, while video carries the raw sequence: motion, sound, objects, speech, context, and timing.
TwelveLabs says the funding will help it advance Marengo and Pegasus, its core video models, and scale what it calls a Video Cognition System. In the company’s framing, that system is meant to make video archives searchable at the level of specific seconds, not just filenames, folders, captions, or transcripts.
The company’s blog describes three technical layers: perception, memory, and reasoning. Marengo maps visual, audio, speech, and on-screen text signals into a searchable representation. Pegasus turns those representations into descriptions, answers, summaries, scene boundaries, entities, temporal segments, and semantic context, according to the company.
That architecture matters because TwelveLabs is not pitching a consumer video generator in this announcement. It is pitching infrastructure for organizations sitting on footage they can’t efficiently query. PYMNTS connected the raise to a broader wave of AI-native software categories, including video generators, AI-native search products, coding assistants, and companion apps.
| TwelveLabs’ claim | What it means in practical terms |
|---|---|
| Video is still “dark matter” to machines | Large archives exist, but much of their content is hard to search semantically |
| Every second should be addressable | Users should be able to locate exact moments, not just files |
| Models should be native to video | The company rejects treating video as captions plus sampled frames |
| Reasoning must span archives | Questions may require comparing events across many clips, not one video |
The funding sharpens the race to make video searchable and useful for AI
This raise turns TwelveLabs’ argument into a capitalized test: can video become a first-class input for AI agents, the way text already has? Lee wrote that “the last decade of AI made text programmable,” while video has not yet had the same shift.
“The world’s video is still mostly dark matter to machines,” Lee said, noting that it sits in places like “archives … drones, and satellites,” mostly still accessed “through filenames, folders, captions, transcripts, and human memory.”
The company says video represents “upwards of 90%” of the world’s data, a figure it used in its Series B press materials. That is a company claim, but it explains the size of the bet: if even a fraction of that footage becomes searchable and usable by AI systems, video search stops being a narrow feature and becomes a workflow layer for enterprises.
The confirmed verticals are specific. TwelveLabs says it has traction in media and entertainment and is moving into the public sector, including work with governments. Its press materials also name advertising, security, sports, and automotive as areas driving demand for its platform.
The strongest counterpoint is that video AI is hard to operationalize. The company itself says brute-force approaches fail in both directions: feeding entire video libraries into a model’s context window would require technology and compute that enterprises could not justify, while turning video into a static database creates structure without intelligence. That is the technical gap TwelveLabs says its Video Cognition System is built to close.
Amazon’s role adds another layer. The supplied materials say Amazon participated in the round, while additional company materials describe AWS as TwelveLabs’ preferred cloud provider and say the models are distributed through Amazon Bedrock and TwelveLabs’ own API. That connects the startup’s product story to infrastructure choices, a pressure point we’ve also covered in Runaway AI Spending Forces a Return to Cloud Controls.
The consumer side is also relevant, but only as context. PYMNTS previously described video generators as part of a new AI app boom, alongside AI companions, conversational search, and prompt-based coding tools. That same shift can be seen in adjacent software categories where AI changes how users discover, trust, and act on information, including the issue raised in Shopify Trustpilot Deal Puts AI-Era Trust on the Line.
Twelve Labs now has to turn video AI hype into enterprise adoption
The raise gives TwelveLabs room to build, but the next proof point is adoption, not vocabulary. Terms like Video Superintelligence and Video Cognition System sound ambitious. Enterprise buyers will care whether the system finds the right moment, answers with evidence, works across messy archives, and does so at a cost that makes sense.
The company’s own roadmap points to that test. It says the money will go toward advancing Marengo and Pegasus, scaling the Video Cognition System into major video archives, and expanding the team. It also says it is hiring researchers, engineers, product builders, and operators.
TwelveLabs has already moved beyond models into applications. Its press materials say the company recently launched Rodeo, its first application-layer product, as part of a push to put the system directly in the hands of creators, operators, and decision-makers without requiring integration work.
The open questions are commercial. The supplied materials do not disclose revenue, customer count, valuation, deployment scale, or pricing. They also do not show independent benchmark results for the latest models. That leaves investors and customers watching for evidence that Twelve Labs video AI can move from impressive search demos to daily production use.
The practical metrics are clear:
- Deployments: Named enterprise rollouts in media, public sector, security, sports, advertising, or automotive.
- Developer uptake: Usage through TwelveLabs’ API and Amazon Bedrock.
- Model reliability: Accuracy across long, noisy, multi-speaker, multi-scene footage.
- Workflow depth: Whether customers treat video search as a core operating tool, not a novelty.
- Infrastructure fit: Whether the AWS relationship helps scale workloads without becoming a cost drag.
If TwelveLabs is right, video becomes one of the next major battlegrounds in AI infrastructure because it carries information text cannot preserve. If it is wrong, the $100 million buys time for a hard lesson: enterprises may want searchable video, but they will only pay for it when the answers are reliable, grounded, and fast enough to replace manual review.
The Bottom Line
- The $100 million raise signals strong investor conviction that video could become a major AI interface.
- Twelve Labs is targeting a hard problem: making vast video archives searchable and understandable at precise moments.
- Backing from Amazon, NEA, NAVER Ventures, and others could help the startup scale its video AI models faster.
Twelve Labs Core Video AI Models
| Model | Role |
|---|---|
| Marengo | Maps visual, audio, speech, and on-screen text signals into a searchable representation. |
| Pegasus | Generates descriptions, answers, summaries, scene boundaries, entities, temporal segments, and semantic context. |
Twelve Labs Series B Funding
Sources
Written by
XOOMAR Insights Team
Research and Editorial Desk
The XOOMAR Insights Team pairs automated research with human editorial judgment. We track hundreds of sources across technology, fintech, trading, SaaS, and cybersecurity, cross-check the facts, and explain what happened, why it matters, and what to watch next. We do not just rewrite headlines. Every article is fact-checked and scored for reliability before it goes live, and we link back to the original sources so you can verify anything yourself.
Explore More Topics
Related Articles
TechnologyAnthropic Revenue Hits $65B as AI Spending Peaks
Anthropic's annualized revenue run rate exploded from $9B to $65B in eight months, as its tech giant partners ramped up massive enterprise deployments of Claude
TechnologyAmazon Bulldozes Rare Books for AI-Fueled Data War
Amazon is buying and then destroying rare books in a Las Vegas warehouse to scan their pages and feed text data into its AI training models, a practice uncovere
TechnologyAnthropic CEO Rejects Doomsayer Label Amid $2T IPO Set-Up
Anthropic CEO Dario Amodei defends his focus on AI risks as a necessary builder of public trust, directly rebutting an investor's claim that his warnings are ba
TechnologyNostalgia’s Last Stand Sells for 72 Dollars
Polaroid's second-gen Go camera, bundled with a free film pack for $71.99, offers the tangible ritual of instant photography for less than the hardware alone.
TechnologyAward-Winning Game Slashes Price to Just $33
Clair Obscur: Expedition 33, the 2025 Game of the Year winner, is now priced at just $33 in a steep discount that signals a fierce fight for consumer spending.
CybersecurityApple Spyware Alerts Swamp Targets In 110 Countries
An unprecedented wave of Apple spyware alerts hit targets across 110 countries, signaling a troubling shift from surgical government surveillance to mass-scale
FintechSynchrony Hijacks ChatGPT Shopping Convos for Credit Offers
Financial giant Synchrony just plugged its credit and offers marketplace directly into ChatGPT, aiming to push promotional financing at users the moment they as
TradingRetail Earnings Unmask a 0.6% Consumer Spending Drop
A surprise 0.6% drop in July retail sales signals that consumers are stressed, while earnings this week from Walmart and others will provide the real-time diagn
TradingGold Soars Past $4,400 As US Dollar Weakening Continues
Gold surged over 1% to $4,422 as a weakening US Dollar, driven by expectations of a more dovish Federal Reserve, ignited a sharp rally in the precious metal.
FintechEbanx Plants Executives in Foreign Markets for Growth
Payments giant Ebanx is decentralizing its leadership, placing senior executives directly in high-growth markets to address local complexity and drive its next
Don't miss the signal
Get our weekly roundup of the stories that matter across tech, fintech, and trading. No noise, just signal.
Free forever. No spam. Unsubscribe anytime.