i
News
News · 2026-10-06

TwelveLabs targets robotics training data with Pegasus 1.6

@neuronium_ai @neuronium_ai

TwelveLabs has released Pegasus 1.6, a video model designed to turn first-person recordings of physical work into timestamped, structured data for robotics teams. The update targets a gap between collecting footage and making it useful for training: identifying what a person does, which hand does it and when each action begins and ends. The model can organize video, but it does not teach a robot how to execute the work.

Cover: TwelveLabs targets robotics training data with Pegasus 1.6

From footage to action labels

TwelveLabs built its earlier video tools for media, sports, entertainment and advertising, where customers needed to search large libraries and describe what they contained. Pegasus applies similar capabilities to a different problem: finding useful examples of physical actions.

Teleoperation, in which a person remotely controls a robot, produces valuable training examples. But, TwelveLabs co-founder and CEO Jae Lee told VentureBeat before the announcement, collecting them requires equipment and trained operators. He described the process as one person working with one robot.

First-person recordings offer another source: a camera worn by a worker can capture tasks without deploying a robot for each recording. The footage still needs interpretation. A system must identify the action and object, determine which hand is involved, and mark the action’s boundaries.

Pegasus 1.5 already segmented video with timestamps and produced structured descriptions. TwelveLabs said earlier versions struggled with first-person recordings. Pegasus 1.6 is specifically trained for that footage; the change is the specialization and claimed improvement in quality, not the basic ability to turn video into structured text.

Customers specify the categories and fields they want, and the model organizes video accordingly, Lee said. For an assembly task, a team could request separate segments for grasping, positioning and fastening a component, plus descriptions of each hand’s role. That is an example of a possible workflow, not a disclosed customer deployment.

TwelveLabs lists five intended uses:

Segmenting recordings into labeled actions.
Creating detailed video captions.
Checking recording quality.
Finding unusual events and duplicate footage.
Detecting potentially sensitive material before further processing.

Quality checks can flag obstructed views, unstable cameras and actions that are hard to understand. Pegasus 1.6 also adds built-in image analysis and improved recognition of people and objects, allowing customers to process still images through the same API as video.

The model provides descriptions and structure. Robotics teams still have to connect those observations to movement, sensor data and control systems before a robot can perform a task reliably.

What the evidence does not show

Lee described Pegasus’s near-term role in robotics as training infrastructure. Companies are still actively training models, he said, while current deployments tend to focus on specific tasks or field-data collection to improve models.

His longer-term thesis is that video of people could become a broad source of knowledge about behavior, with teleoperation examples helping adapt that knowledge to real robots. That is a strategy, not evidence that adding more video by itself produces a general-purpose robot.

Pegasus can describe limb movements and roughly estimate trajectories, Lee said, but that is not enough to execute actions. Pressure, touch and precise control remain separate problems; teams must combine video with other data and build systems that translate it into action.

TwelveLabs’ technical appendix describes internal evaluations of action recognition, identifying which hand performs an action and understanding action sequences over time. It provides no numerical results or reproducible comparison with competitors. The company also did not name a robotics customer whose deployment results could be independently checked.

Lee recalled one processing run that handled about 17 years of first-person video in roughly 18 hours, with no video failing. He did not specify the compute resources or evaluation conditions. That is a claim about processing speed, not a reproducible benchmark or measure of accuracy.

I think the missing comparison matters more than the processing anecdote. Teams need to know not just how much footage the model can process, but how often its labels match expert judgment and how much human correction is needed before those labels are useful.

A place in the robotics stack

The competing products address different stages of robotics development. Nvidia’s Cosmos Curator offers open-source tools to filter, label and deduplicate data. Encord combines Nvidia Cosmos Reason 2 and Embed with hosted models, human review and action-based search. Google’s Gemini Robotics ER 2 analyzes video and tracks task progress, with a greater focus on planning and coordination; lower-level models or robotics interfaces handle movement.

FLUX-mimic, from Black Forest Labs and mimic, uses a video-model foundation to generate robot actions. The companies report trials and deployment at Audi, including industrial manipulation. It is an example of a system that could use better training data, but no integration with TwelveLabs was disclosed.

Pegasus 1.6 is priced at $1.75 per hour of video, $3 per million input image tokens and $7.50 per million output tokens, roughly the same output-token price as Pegasus 1.5. Enterprise terms are negotiated individually. At the listed video rate, processing 1,000 hours once costs $1,750 before output tokens and other possible charges. More detailed labels can increase output costs, and TwelveLabs charges for each segment description requested in a Segment call.

Those rates do not establish which option is cheapest for a particular task. Pegasus bills video by duration and text output by tokens; Google bills input and output tokens. Open-source software also requires compute, storage and engineering work. The available information is a feature comparison, not a ranking of quality, cost or downstream robotics results.

My guess is that the commercial test is not whether Pegasus can describe a task, but whether it can produce accepted training examples for less than a team can assemble and maintain its own pipeline. The company is selling a step in the data workflow, not robot behavior itself.

A useful pilot would focus on one task, such as packing a specific product or assembling a particular component. Teams can compare the model’s labels with expert review, count the corrections required, and test whether the resulting data improves performance on examples not used for training. They also need to decide how worker recordings may be collected and how data about internal processes will be handled; detecting sensitive images does not itself grant permission to record or use them.

The decisive metric is the cost of an accepted, training-ready example—and whether that example improves a robot’s behavior. Pegasus can shorten the path from a recording to labeled data. Whether that saves work depends on the accuracy and usefulness of the labels, not the volume of footage processed.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X