Google DeepMind 联合哈佛、斯坦福等实验室发表论文,提出通往 AGI 的路径可能是构建世界模型的视觉 AI——能记忆变化、预测结果并采取行动。论文主张视觉系统应直接从图像、视频、3D 结构和交互中学习,理解存在什么、变化了什么、隐藏了什么、接下来可能发生什么,而非仅将视觉输入语言模型。
Google DeepMind + Harvard + Stanford and many other top labs paper argues that a path to AGI may be visual AI that builds a world model, remembers changes, predicts outcomes, and acts.
Most multimodal AI still treats vision as something you feed into a language model.
The paper wants vision to do more of the thinking itself.
A capable visual system should learn directly from images, video, 3D structure, and interaction. It should understand what exists, what changed, what is hidden, what might happen next, and what it needs to look at before acting.
they point to video generation, reconstruction, persistent memory, continual learning, multimodal sensing, and robotics as possible pieces of the same system.
they say stop judging visual AI mainly by image Q&A, captions, or realistic video.
learn the world's structure from visual experience, keep updating that knowledge, and use it to predict and act.
来源:@rohanpaul_ai · x.com