i
DATAIST
News · 2026-09-20

Runway targets live, controllable video generation

@neuronium_ai @neuronium_ai

Runway is preparing a video-generation system that behaves more like a live feed than a rendering queue. Instead of waiting seconds or minutes for a finished clip, users would see the first frame quickly and receive new frames as they change the prompt. The shift matters because it moves generative video from one-shot production toward continuous control.

Cover: Runway targets live, controllable video generation

From finished clips to live control

Most video models work in stages: prompt first, result later. If the output misses the mark, the user starts again. Runway says users spend most of their time creating and correcting video, so its first goal is to reduce the delay before the first frame, then stream the result as the prompt changes.

The company first described this direction in March with Runway Characters. The system is built on GWM-1, Runway’s first “universal world model,” introduced in December 2025.

GWM-1 extends Gen-4.5 and generates video frame by frame. It can take several kinds of control signals:

Camera movement
Commands to a robot
Sound

A few weeks ago, Runway also showed Solaris, a system based on Gen-4.5 that generates user interfaces frame by frame and responds to clicks and voice input.

The common thread is not simply faster rendering. It is a change in the interface: video becomes something the user steers while it is being created.

Source: the-decoder.com

The cost of correcting a frame

Runway argues that real-time generation could narrow the distance between an idea and its execution. With immediate feedback, users would spend more time directing the video and less time waiting for it.

The economics are part of the pitch. Faster models use less GPU time and therefore cost less to run. Runway’s argument is that an application’s viability depends on the price of one result at a given quality level. Instant generation could lower that threshold enough to support applications that previously did not pay for themselves.

The technical problem is error accumulation. A text model can correct itself halfway through a sentence. A video model builds each frame from the previous one, so small mistakes can compound into severe distortions. Runway calls this the central weakness of approaches based on large language models.

Its answer is to train the model not only on clean data, but also on its own outputs. The model then learns to correct deviations instead of amplifying them.

Decart used a similar method in MirageLSD, its real-time model, by deliberately exposing it during training to damaged and distorted images. Google DeepMind says its world model Genie 3 can maintain consistent interactive worlds for several minutes at 24 frames per second and 720p.

Unlike language models, video models can't correct an earlier error because each new frame builds on the flawed one. | Image: Runway

Unlike language models, video models can't correct an earlier error because each new frame builds on the flawed one. | Image: Runway

Source: the-decoder.com

Real-time generation also changes where the computing burden sits. Runway says more of the load moves from training to use: the model has to produce every frame quickly enough to keep pace with playback while running on hardware that serves multiple sessions at once.

That tradeoff is easy to miss in the live-demo framing. A model can feel instantaneous in a tightly controlled example and still be expensive or fragile at the scale of a real application.

Source: the-decoder.com

World models beyond media

Runway sees interactive applications as the long-term direction for AI-generated media. The company points to education, games and robotics as areas where video needs to react at roughly the speed at which people perceive it.

Robotics and autonomous driving add a second use case: simulated environments that can generate unusual situations in real time and respond immediately. Runway previously introduced GWM Robotics, a version of GWM-1 for generating synthetic training data for robots.

Waymo is pursuing a similar idea with a world model based on Genie 3 and adapted for road traffic. It lets the company model situations its vehicle fleet has not yet encountered, including:

An elephant
A tornado
A flooded residential area

Waymo says Waymo Driver travels billions of virtual miles before encountering comparable scenarios on real roads.

In March, Runway also showed a research version of a real-time model developed with Nvidia at the chipmaker’s GTC conference. It runs on the Vera Rubin platform and is designed to produce its first frame in less than 100 milliseconds.

under 100 millisecondsfirst frame
24frames per second
720presolution

My read is that Runway is defining video generation less as a content pipeline than as a control layer for simulated worlds. That makes the robotics and autonomous-driving examples more important than the promise of faster clips: the value comes from responding to an action, not merely producing a polished output.

What I’d want to know is what Runway is willing to give up to achieve that responsiveness. The company has not announced when the research model will be available, and the announcement does not establish how consistency, quality and cost change when several sessions are running at once. Until those constraints are visible, “real time” describes a direction more clearly than a product.

The strategic tension is straightforward: the closer generated video gets to a live control loop, the less it can be judged only as a finished image—and the more its usefulness depends on staying coherent under pressure.

Source: the-decoder.com

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X