Urgent.News

What's breaking now, across thousands of outlets.

AI

Beyond LLMs: How World Models Are Changing Generative Media

If I give my daughter milk or water, I can almost guarantee she will spit it out. She’s 21 months old, and while she actually enjoys both drinks, halfway through she inevitably decides it’s time for a daily physics experiment. She wants to know: What will happen if I open my mouth and just let the water fall out? How Humans Learn the World While it annoys me, I remind myself that this is how…

Ever since my 21-month-old daughter started spitting out milk or water halfway through her drink, I’ve come to realize that this behavior mirrors how humans learn about the world. By observing cause-and-effect through repetition, we build a mental map, allowing us to predict outcomes without constantly testing them. This concept is referred to as an internal world model, a mental representation of how reality operates that aids in anticipating potential consequences before taking action.

While Large Language Models (LLMs) have dominated the generative AI landscape over the past few years, a more ancient notion in AI is making a resurgence—world models. These systems learn from environmental patterns and how they change over time, including the outcomes of actions within that environment. Much like predicting the spillage of water when opening a full glass, world models can infer what will happen next based on learned patterns.

Historically, world models were primarily employed in robotics and autonomous vehicles. Autonomous cars, for example, rely on world models to navigate safely: if the car changes lanes or brakes suddenly, it must predict the resulting actions to prevent accidents. The application of world models is expanding beyond these domains, however, with companies like Runway incorporating them into video generation and user interfaces.

Runway’s GWM Worlds 2 is a prime example, turning video generation into a real-time interactive simulation. Instead of waiting for a finished video after inputting a prompt, users can continue steering the environment while it’s being generated. They define the world’s parameters—environment, subjects, style, and physical rules—and then manipulate it through text or camera movement to influence the outcome.

For instance, asking the model to "Let's make it rain in this video" would alter the scene accordingly, changing how characters, buildings, and other elements respond.

More intriguingly, world models could revolutionize software interfaces. Runway recently launched Solaris, an Interface World Model that dynamically generates interface elements in response to user interactions. In a traditional banking app with fixed tabs, a world model could allow users to ask about their spending and transfer money, with the interface generating charts, explanations, and controls on the fly. This approach could streamline user experiences, eliminating the need to navigate predetermined screens.

While such advancements in generative media hold significant promise across tech, media, film, and gaming, my immediate interest lies in interactive spatial design. Imagine browsing homes on Zillow and visualizing how you would furnish a room that's beyond your budget. With a world model, you could prompt questions like "Can a king-sized bed fit here?" or "How would morning sunlight hit this wall?"

The environment would then adapt, allowing you to modify elements like the bed's position, wall color, or partition removal. This interactive approach could transform not just real estate and interior design, but also the broader trajectory of generative media.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 10 September →