Home›AI Comparisons›Runway’s GWM Worlds 2: Real-Time Interactive World Models Explained
AI Comparisons

Runway’s GWM Worlds 2: Real-Time Interactive World Models Explained

Runway’s GWM Worlds 2 introduces real-time interactive video and audio generation using WorldPrompt, a new prompting system that enables dynamic, timestamped actions in simulated environments. This marks a significant step toward self-generated, interactive AI worlds.

Real-time interactive world simulation interface with animated characters and action timeline

Runway, the AI-powered creative platform valued at $5.3 billion, has released GWM Worlds 2, a research preview that enables real-time interactive video and audio generation. This advancement builds on its earlier work with generative video models and introduces WorldPrompt, a novel prompting mechanism designed to steer a world model through time with persistent context and timed actions.

What Happened: A New Layer of Control in AI Worlds

Launched in September 2026, GWM Worlds 2 represents a shift from static video generation to dynamic, interactive simulation. Unlike traditional AI video models that produce a complete clip at once, GWM Worlds 2 generates video and audio in real time—streaming at 720p resolution and 24 frames per second, with audio at 48,000 Hz.

The centerpiece of this release is WorldPrompt, a control layer that allows users to define a simulated environment and specify a sequence of timestamped actions. For instance, a user can fix the first frame of a scene and then define events such as an NPC walking toward the viewer or a camera panning to a new location. These actions can even be prompted in real time, enabling responsive interactions.

According to Runway’s CTO Kamil Sindi and Principal Research Scientist Robin Kahlow, WorldPrompt is not a programming language or scriptable environment like Minecraft or Roblox. Instead, it functions as a prompt-based interface that enables detailed control over characters, cameras, and environmental elements—without requiring full state management or scripting capabilities.

Key Facts and Technical Foundations

  • GWM Worlds 2 uses an autoregressive diffusion model, which generates video frame by frame rather than producing a full clip at once.
  • The model was fine-tuned to interpret the WorldPrompt format and then post-trained to generate content in a real-time, autoregressive manner.
  • Runway achieved real-time performance through distillation techniques—reducing the number of denoising steps from around 50 to four—while maintaining acceptable visual and audio quality.
  • WorldPrompt enables real-time, interactive actions, such as movement or dialogue, with reliable performance for basic behaviors like walking or speaking.
  • Despite progress, the model does not have perfect long-term memory, and error accumulation remains a significant technical challenge.

The system is designed to produce plausible outcomes under different user actions—a key distinction between current video models and true world models. As co-CEO Anastasis Germanidis explained in a podcast, video models trained on real-world data often favor successful outcomes (e.g., successful football goals), leading to biased or unrealistic counterfactuals. A world model must instead generate equally realistic outcomes for both success and failure.

How It Works: From Video Generation to Real-Time Simulation

GWM Worlds 2 operates through a multi-stage engineering process. First, Runway fine-tuned its foundational audio-video generation model to interpret the WorldPrompt format. Then, it post-trained the model to generate content in an autoregressive fashion—producing one frame at a time, feeding the output back into the model to generate the next.

This autoregressive approach avoids the computational burden of generating a full video in advance. However, it introduces the risk of error accumulation: small inaccuracies in one frame can compound over time, leading to degraded realism in longer sequences.

To mitigate this, Runway employs distillation methods—either reducing the size of the original model or minimizing the number of diffusion steps. For example, a model may transition from 50 denoising steps to four, with a trade-off in fidelity but improved speed and real-time responsiveness.

Another key challenge is managing memory. The model must continuously maintain relevant context while discarding outdated or irrelevant information to prevent excessive GPU memory usage. This requires sophisticated optimization strategies, especially in long-duration simulations.

Why It Matters: A Step Toward Self-Generated Interactive Worlds

While current systems like GWM Worlds 2 are still in a research preview phase, they represent a foundational step toward fully self-generated, interactive AI worlds. These environments could eventually power AI-driven games, virtual training simulations, or immersive educational experiences.

500px provided description: A close up view of a glowing digital city that is made of just buttons and knobs. [#blue ,#glow ,#studio ,#green ,#music ,#tv ,#florida ,#digital ,#close up ,#glowing ,#tech ,#buttons ,#production ,#electronic ,#volume ,#blog ,#radio ,#mixer ,#levels ,#knobs ,#akai ,#sliders ,#nikon d800 ,#dance music ,#headers ,#soundcloud ,#apc40 ,#digital city ,#dj tools ,#producers tools]
500px provided description: A close up view of a glowing digital city that is made of just buttons and knobs. [#blue ,#glow ,#studio ,#green ,#music ,#tv ,#florida ,#digital ,#close up ,#glowing ,#tech ,#buttons ,#production ,#electronic ,#volume ,#blog ,#radio ,#mixer ,#levels ,#knobs ,#akai ,#sliders ,#nikon d800 ,#dance music ,#headers ,#soundcloud ,#apc40 ,#digital city ,#dj tools ,#producers tools] by Greg Sagayadoro, CC BY-SA 3.0, via Wikimedia Commons. · Source · License

As Robin Kahlow emphasized, the core challenge lies in making video generation both fast and causally consistent. A world model must not only generate frames in real time but also produce plausible, logically consistent outcomes when users take different actions—such as choosing to open a door or throw a ball.

Runway’s work aligns with broader trends in generative AI, where models are moving from static outputs to dynamic, interactive systems. This shift echoes the evolution of creative tools—from early image segmentation and rotoscoping to modern diffusion models—showing how AI tools have matured from niche applications to production-grade capabilities.

For the broader AI ecosystem, GWM Worlds 2 illustrates how real-time world modeling is no longer a distant vision. Instead, it is being engineered through a combination of model architecture, prompt design, and real-time optimization.

Limitations and Open Questions

Despite its promise, GWM Worlds 2 has notable limitations. The model does not have perfect long-term memory, meaning it may fail to retain or correctly respond to complex, multi-step instructions. For instance, it cannot reliably implement abstract rules—like gravity or character abilities—without explicit, simple prompts.

Additionally, error accumulation poses a persistent risk. As each frame is generated based on the previous one, small inaccuracies can grow over time, degrading the realism of longer sequences. This issue is particularly pronounced in scenarios involving complex interactions or physical laws.

Another open question is how well the model handles counterfactual generation—producing realistic outcomes for actions that are not observed in training data. For example, simulating a failed football shot should be as plausible as a successful one, yet current models may still favor the former due to biased training data.

What to Watch Next

As GWM Worlds 2 evolves, several developments will be critical to track. First, improvements in long-term memory and error mitigation will determine how reliably the model can maintain consistent world states over time. Second, progress in counterfactual generation will reveal whether these models can simulate realistic, diverse outcomes under varied user inputs.

Runway’s trajectory—moving from creative tools to real-world simulation—mirrors the broader evolution of generative AI. As the company continues to refine its models, it may eventually integrate with other AI systems, such as autonomous agents or real-time simulation environments.

For readers interested in the broader AI infrastructure, The Future of Latent Space offers deeper insight into how foundational AI research is shaping current tools. Similarly, OpenRouter’s journey provides context on how AI infrastructure scales from research to enterprise use.

For a comparative view of real-time world models, see Google DeepMind’s Genie 3 and Runway’s official GWM Worlds 2 page.

Sources & further reading

Featured image: Carola is interviewed by Gina Dirawi in the Swedish TV-show Studio Eurovision in 2013. by Albin Olsson, CC BY-SA 3.0, via Wikimedia Commons. Image source · License

SIRIUSITY BRIEFS

Get the latest from Siriusity Journal

AI, science, technology and future-tech updates delivered to your inbox.

We use your email only for these briefs. Double opt-in is required.