Urgent.News

What's breaking now, across thousands of outlets.

Tech

Atlas: A World Model for Spatial Intelligence

Article URL: https://www.worldlabs.ai/blog/atlas Comments URL: https://news.ycombinator.com/item?id=49525160 Points: 204 # Comments: 49

World models are versatile tools that generate, reconstruct, and simulate various possible worlds. At World Labs, researchers are developing advanced world models with the goal of achieving spatial intelligence. Today, they unveil Atlas, a next-generation world model designed to work natively with text, images, video, and 3D data.

Atlas utilizes a multimodal autoregressive diffusion transformer architecture, where all inputs are integrated into a shared spatial context. This context allows Atlas to generate consistent 3D outputs and envision future elements based on the information it has encountered.

The model is built for scalability, with its performance improving as training compute increases. This scalability trend is expected to continue as the model is further scaled. Atlas excels at various tasks, including world generation, reconstruction, and simulation. It can produce new views from reference images at any desired camera position and angle, smoothly extrapolating beyond the input to create unseen scene elements.

With precise camera geometry as a native input, Atlas offers greater control over camera framing and motion compared to traditional text-based instructions.

One of Atlas's unique features is its ability to incorporate precise 3D camera geometry, enabling it to generate worlds based on spatial context. This capability opens up new creative possibilities, such as smoothly interpolating between unrelated image pairs to create coherent worlds. Atlas can generate extensive videos with precise control, allowing users to design every scene and camera angle. In this way, users become directors rather than relying on random generation.

Atlas also excels at reconstructing real-world spaces from limited input images. It can generate faithful reconstructions with as few as two or three images, outperforming state-of-the-art models that specialize solely in 3D reconstruction. When more input images are provided, Atlas refines its reconstruction, eliminating the need for extensive capture equipment or dense views.

This capability is particularly useful for generating aerial views from ground-level photos and piecing together complex scenes from multiple images.

Additionally, Atlas can generate various camera trajectories through the same scene, offering new perspectives on the input images. This flexibility enables users to experiment with different viewpoints, moods, and visual styles. Furthermore, Atlas can output worlds in both 2D image frames and 3D depth maps, making it suitable for a wide range of applications in robotics, gaming, design, VFX, and beyond.

Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at worldlabs.ai →

More in Tech

What’s the Scam?

To subscribe to my monthly email newsletter, you have to enter your information on the webpage, and then reply to an automatically generated email.

More from Tuesday 1 September →