World Labs, the AI startup co-founded by computer vision pioneer Fei-Fei Li, has unveiled Atlas, a “world model” that builds explorable 3D scenes from a single image. The release is the fullest expression yet of Li’s theory of spatial intelligence, the idea that AI must reason natively about objects, people and environments to understand the physical world.
Under the hood sits what World Labs calls an omni model: a multimodal autoregressive diffusion transformer it claims outstrips first-generation video tools. Rather than steering shots with fuzzy text prompts, Atlas treats camera paths and geometry as first-class inputs, so creators dictate exactly where the lens goes.
A single 2D photograph is enough to produce as much as a minute of 1440p footage whose geometry holds from any viewing angle. The same pipeline yields 3D assets, from point clouds to Gaussian splats, alongside video, depth maps and camera poses.
World Labs points to creative effects and game design as early uses, but its real target is robotics. Developers can scan a physical space with an ordinary smartphone, reconstruct it as a simulation, then let Atlas render photorealistic images and depth readings as a robot moves through it. That gives hardware teams a cheap, safe way to test across lighting conditions and object arrangements.
The startup has raised $1.2B from backers including Nvidia, AMD and Autodesk since its February 2024 founding.