AI 导读
World Labs 发布 Atlas,一个从零训练的多模态世界模型,能生成具有像素级相机控制的画面、仅凭一张输入图像重建大场景、通过重构图视频模拟时空,并从一个或多个输入图像原生输出 3D 空间,还能把多张带位姿图像合成为一致的 3D 世界。该模型被描述为相机条件世界模型,可用于 VFX 到机器人等多种场景。转发者评论称技术路线从 LLM 走向 World Model,再到 Physical AI。
正文
从 LLM 到 World Model,再到 Physical AI。
先理解你,再理解世界,最后让机器人出来统治你。
以上。 https://t.co/ozqKcm70dI
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀 Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️在 X 查看被引用的帖子
来源:@frxiaobei · x.com