World Labs 发布多模态世界模型 Atlas,称其从零训练,可生成像素级相机控制的画面、从单张图像重建大场景、通过重构图视频模拟时空,并能从一张或多张图像原生输出 3D 空间。
Why is Spatial Intelligence considered the lifeline for Embodied AI?
Because the biggest Achilles' heel preventing robots from being deployed in the real world is the Sim2Real (Simulation to Reality) gap and a lack of generalization:
Setting up 1,000 rooms in the real world for destructive testing would bankrupt any startup;
whereas a world model like Atlas gives AI the ability to autonomously generate an infinite number of hyper-realistic physical environments from limited real-world data.
This is the true ambition behind Fei-Fei Li founding World Labs—not to build just another video generator, but to create the digital foundation of the physical world. https://t.co/3rI1YhcUd5
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀 Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️在 X 查看被引用的帖子
来源:@AYi_AInotes · x.com