Overview
Atlas is the world's first multimodal world model released on September 1, 2026, by World Labs, a spatial intelligence company co-founded by Li Feifei. The model is pre-trained from scratch and natively processes text, images, videos, camera poses, and 3D depth information. Through a spatial context mechanism, it anchors various inputs to unified 3D spatial coordinates, enabling generation based on spatial reasoning. The release of Atlas is regarded as a significant milestone in the spatial intelligence field, often referred to as the 'GPT moment' for spatial intelligence. Currently, early access is only available to partners, with no public pricing or registration.
Key Features
- Camera-Controlled Generation: Takes camera pose parameters as native input. Given 1-6 ordinary photos and camera movement trajectories, it can generate videos up to 1 minute long at 1440p resolution, achieving pixel-level precise camera control.
- Spatial Reconstruction: Reconstructs real scenes from 2-25 tourist-style flat photos, outputting explicit 3D representations such as point clouds and 3D Gaussian splats. On public datasets like DTU, ETH3D, and ScanNet, the sparse-view reconstruction error (AbsRel) is 25.3.
- Spatiotemporal Simulation: Supports bullet-time effects, allowing regeneration of scenes from any angle using 3-5 camera positions. It also supports video re-composition and can convert real videos into simulation environments for robot training, generating body-camera RGB and depth data.
- Image Generation: Supports text-to-image and 360-degree panorama generation, capable of handling complex prompts and rendering text within images.
- Multimodal Unified Architecture: Employs a multimodal autoregressive diffusion Transformer. The spatial context mechanism anchors each input image to specific coordinates in 3D space, allowing the model to reason in 3D space rather than pixel sequences. It can stitch two unrelated reference images and set positions to generate transitional worlds.
Use Cases
- Visual effects production: Quickly reconstruct spatial scenes from a few on-site photos for post-production and virtual production.
- Game design: Generate controllable-view 3D environments and dynamic scenes to assist level design and concept validation.
- Robot training: Convert real operation videos into simulation environment data, generating body-camera RGB and depth information to reduce data collection costs.
- Spatial content creation: Generate 360-degree panoramas and multi-view videos from text or sparse photos for immersive experience production.
Pros
- Native support for camera pose input, achieving pixel-level precision in video generation camera control.
- Outstanding sparse-view reconstruction capability, outputting explicit 3D structures from just a few flat photos.
- Multimodal unified architecture design, sharing the same spatial reasoning framework across text, images, videos, and 3D depth information.
- Supports Real-to-Sim conversion for robots, providing an efficient data pathway for simulation training.
- Can stitch multiple reference images and specify 3D positions to generate transitional worlds, offering high creative flexibility.
- Led by Li Feifei's team, with investments from NVIDIA, AMD, and other institutions, backed by solid technical expertise and resources.
Pricing
Currently, early access is only available to partners, with no public pricing or registration channels. Specific commercial terms are subject to official announcements.
Summary
As the world's first multimodal world model, Atlas unifies text, images, videos, camera poses, and 3D depth in spatial context for reasoning, demonstrating comprehensive capabilities in camera-controlled generation, sparse reconstruction, and spatiotemporal simulation. Its applications span visual effects, game design, and robot training data infrastructure. Currently in the early access phase for partners, it will drive the development of products like Marble under World Labs in the future.
Version History
- World Labs Atlas official launch (2026-09-01): World Labs unveiled Atlas, billed as the world's first multimodal world model: a spatial-context mechanism anchors text, images, video, camera poses and 3D depth in 3D space, enabling pixel-accurate camera-controlled video generation (up to 1 minute at 1440p from 1-6 photos plus a camera path), explicit 3D reconstruction (point clouds and 3D Gaussian splats) from sparse photos, space-time simulation and Real-to-Sim robotics training. It is currently in partner early access with no public pricing, and will later power products such as Marble.