XOOMAR
Futuristic lab with video frames and neural networks symbolizing scaled video AI development
TechnologyJuly 1, 2026· 6 min read· By XOOMAR Insights Team

Twelve Labs Grabs $100M as Video AI Battles Chatbots

Share
Updated on July 2, 2026

Twelve Labs video AI just pulled in $100 million because investors are making AI startup bets that the next AI interface won’t be a chat box, it will be searchable footage. The startup said Wednesday, July 1, that it raised a $100 million Series B to expand its work on models that can understand, index, retrieve, and reason over video, according to PYMNTS.

XOOMAR Intelligence

Analyst Take

72/ 100
High
4 sources analyzedMedium confidenceTrend10Freshness96Source Trust88Factual Grounding94Signal Cluster20

The round was co-led by NEA and NAVER Ventures, with participation from Amazon, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital, and Red Bull Ventures, according to TwelveLabs’ own announcement. The company did not disclose a valuation in the supplied materials.

“Five years ago, we began with a simple observation: The world does not happen in text. It happens in motion,” Co-Founder and CEO Jae Lee wrote.

Twelve Labs lands $100 million to scale its video AI platform

The core thesis behind Twelve Labs video AI is blunt: text has been made programmable, but video remains largely locked away from machines. Lee argues that most AI systems still work from compressed descriptions of reality, while video carries the raw sequence: motion, sound, objects, speech, context, and timing.

TwelveLabs says the funding will help it advance Marengo and Pegasus, its core video models, and scale what it calls a Video Cognition System. In the company’s framing, that system is meant to make video archives searchable at the level of specific seconds, not just filenames, folders, captions, or transcripts.

The company’s blog describes three technical layers: perception, memory, and reasoning. Marengo maps visual, audio, speech, and on-screen text signals into a searchable representation. Pegasus turns those representations into descriptions, answers, summaries, scene boundaries, entities, temporal segments, and semantic context, according to the company.

That architecture matters because TwelveLabs is not pitching a consumer video generator in this announcement. It is pitching infrastructure for organizations sitting on footage they can’t efficiently query. PYMNTS connected the raise to a broader wave of AI-native software categories, including video generators, AI-native search products, coding assistants, and companion apps.

TwelveLabs’ claim What it means in practical terms
Video is still “dark matter” to machines Large archives exist, but much of their content is hard to search semantically
Every second should be addressable Users should be able to locate exact moments, not just files
Models should be native to video The company rejects treating video as captions plus sampled frames
Reasoning must span archives Questions may require comparing events across many clips, not one video

The funding sharpens the race to make video searchable and useful for AI

This raise turns TwelveLabs’ argument into a capitalized test: can video become a first-class input for AI agents, the way text already has? Lee wrote that “the last decade of AI made text programmable,” while video has not yet had the same shift.

“The world’s video is still mostly dark matter to machines,” Lee said, noting that it sits in places like “archives … drones, and satellites,” mostly still accessed “through filenames, folders, captions, transcripts, and human memory.”

The company says video represents “upwards of 90%” of the world’s data, a figure it used in its Series B press materials. That is a company claim, but it explains the size of the bet: if even a fraction of that footage becomes searchable and usable by AI systems, video search stops being a narrow feature and becomes a workflow layer for enterprises.

The confirmed verticals are specific. TwelveLabs says it has traction in media and entertainment and is moving into the public sector, including work with governments. Its press materials also name advertising, security, sports, and automotive as areas driving demand for its platform.

The strongest counterpoint is that video AI is hard to operationalize. The company itself says brute-force approaches fail in both directions: feeding entire video libraries into a model’s context window would require technology and compute that enterprises could not justify, while turning video into a static database creates structure without intelligence. That is the technical gap TwelveLabs says its Video Cognition System is built to close.

Amazon’s role adds another layer. The supplied materials say Amazon participated in the round, while additional company materials describe AWS as TwelveLabs’ preferred cloud provider and say the models are distributed through Amazon Bedrock and TwelveLabs’ own API. That connects the startup’s product story to infrastructure choices, a pressure point we’ve also covered in Runaway AI Spending Forces a Return to Cloud Controls.

The consumer side is also relevant, but only as context. PYMNTS previously described video generators as part of a new AI app boom, alongside AI companions, conversational search, and prompt-based coding tools. That same shift can be seen in adjacent software categories where AI changes how users discover, trust, and act on information, including the issue raised in Shopify Trustpilot Deal Puts AI-Era Trust on the Line.


Twelve Labs now has to turn video AI hype into enterprise adoption

The raise gives TwelveLabs room to build, but the next proof point is adoption, not vocabulary. Terms like Video Superintelligence and Video Cognition System sound ambitious. Enterprise buyers will care whether the system finds the right moment, answers with evidence, works across messy archives, and does so at a cost that makes sense.

The company’s own roadmap points to that test. It says the money will go toward advancing Marengo and Pegasus, scaling the Video Cognition System into major video archives, and expanding the team. It also says it is hiring researchers, engineers, product builders, and operators.

TwelveLabs has already moved beyond models into applications. Its press materials say the company recently launched Rodeo, its first application-layer product, as part of a push to put the system directly in the hands of creators, operators, and decision-makers without requiring integration work.

The open questions are commercial. The supplied materials do not disclose revenue, customer count, valuation, deployment scale, or pricing. They also do not show independent benchmark results for the latest models. That leaves investors and customers watching for evidence that Twelve Labs video AI can move from impressive search demos to daily production use.

The practical metrics are clear:

  • Deployments: Named enterprise rollouts in media, public sector, security, sports, advertising, or automotive.
  • Developer uptake: Usage through TwelveLabs’ API and Amazon Bedrock.
  • Model reliability: Accuracy across long, noisy, multi-speaker, multi-scene footage.
  • Workflow depth: Whether customers treat video search as a core operating tool, not a novelty.
  • Infrastructure fit: Whether the AWS relationship helps scale workloads without becoming a cost drag.

If TwelveLabs is right, video becomes one of the next major battlegrounds in AI infrastructure because it carries information text cannot preserve. If it is wrong, the $100 million buys time for a hard lesson: enterprises may want searchable video, but they will only pay for it when the answers are reliable, grounded, and fast enough to replace manual review.

The Bottom Line

  • The $100 million raise signals strong investor conviction that video could become a major AI interface.
  • Twelve Labs is targeting a hard problem: making vast video archives searchable and understandable at precise moments.
  • Backing from Amazon, NEA, NAVER Ventures, and others could help the startup scale its video AI models faster.

Twelve Labs Core Video AI Models

ModelRole
MarengoMaps visual, audio, speech, and on-screen text signals into a searchable representation.
PegasusGenerates descriptions, answers, summaries, scene boundaries, entities, temporal segments, and semantic context.

Twelve Labs Series B Funding

Series B
$ million100
XOOMAR

Written by

XOOMAR Insights Team

Research and Editorial Desk

The XOOMAR Insights Team pairs automated research with human editorial judgment. We track hundreds of sources across technology, fintech, trading, SaaS, and cybersecurity, cross-check the facts, and explain what happened, why it matters, and what to watch next. We do not just rewrite headlines. Every article is fact-checked and scored for reliability before it goes live, and we link back to the original sources so you can verify anything yourself.

Related Articles

Modern workspace showcasing financial analysis with digital charts and reports.Technology

Anthropic Revenue Hits $65B as AI Spending Peaks

Anthropic's annualized revenue run rate exploded from $9B to $65B in eight months, as its tech giant partners ramped up massive enterprise deployments of Claude

Aug 18, 20266 min
Close-up image of a laptop keyboard illuminated with blue light, showcasing modern technology design.Technology

Amazon Bulldozes Rare Books for AI-Fueled Data War

Amazon is buying and then destroying rare books in a Las Vegas warehouse to scan their pages and feed text data into its AI training models, a practice uncovere

Aug 17, 20267 min
A woman in VR gear surrounded by computers and cables, immersed in a virtual simulation.Technology

Anthropic CEO Rejects Doomsayer Label Amid $2T IPO Set-Up

Anthropic CEO Dario Amodei defends his focus on AI risks as a necessary builder of public trust, directly rebutting an investor's claim that his warnings are ba

Aug 16, 20266 min
From above instant photos of cute dog placed on desk near camera and smartphoneTechnology

Nostalgia’s Last Stand Sells for 72 Dollars

Polaroid's second-gen Go camera, bundled with a free film pack for $71.99, offers the tangible ritual of instant photography for less than the hardware alone.

Aug 16, 20266 min
Black and white image of a classic Apple II computer on display in Wrocław, Poland.Technology

Award-Winning Game Slashes Price to Just $33

Clair Obscur: Expedition 33, the 2025 Game of the Year winner, is now priced at just $33 in a steep discount that signals a fierce fight for consumer spending.

Aug 15, 20266 min
Top view of a smartphone showing activation lock screen on light blue surface.Cybersecurity

Apple Spyware Alerts Swamp Targets In 110 Countries

An unprecedented wave of Apple spyware alerts hit targets across 110 countries, signaling a troubling shift from surgical government surveillance to mass-scale

Aug 17, 20266 min
Hand holding smartphone displaying digital wallet app interface, blurred monitor in background.Fintech

Synchrony Hijacks ChatGPT Shopping Convos for Credit Offers

Financial giant Synchrony just plugged its credit and offers marketplace directly into ChatGPT, aiming to push promotional financing at users the moment they as

Aug 17, 20265 min
Detailed financial trading screen with colorful charts and data representing market fluctuations.Trading

Retail Earnings Unmask a 0.6% Consumer Spending Drop

A surprise 0.6% drop in July retail sales signals that consumers are stressed, while earnings this week from Walmart and others will provide the real-time diagn

Aug 17, 20267 min
Bitcoin coin on a tablet showing stock chart, surrounded by dollar bills.Trading

Gold Soars Past $4,400 As US Dollar Weakening Continues

Gold surged over 1% to $4,422 as a weakening US Dollar, driven by expectations of a more dovish Federal Reserve, ignited a sharp rally in the precious metal.

Aug 18, 20267 min
A smartphone displaying an ecommerce site with a credit card, set on a wooden surface, depicting online shopping.Fintech

Ebanx Plants Executives in Foreign Markets for Growth

Payments giant Ebanx is decentralizing its leadership, placing senior executives directly in high-growth markets to address local complexity and drive its next

Aug 18, 20266 min

Don't miss the signal

Get our weekly roundup of the stories that matter across tech, fintech, and trading. No noise, just signal.

Free forever. No spam. Unsubscribe anytime.