AI 导读
Cosmos 3 被原生训练用于生成动作,其输入输出覆盖文本、图像、视频、音频和动作。同一份 checkpoint 可分别作为视觉语言模型、视频世界模型或机器人策略运行,无需多模型编排。
正文
3/ Inputs and outputs span text, image, video, audio AND action.
That last one is the big deal. Cosmos 3 was trained natively to generate actions, so the same checkpoint can run as a vision-language model, a video world model, or a robot policy. No multi-model orchestration.
Video
来源:@kimmonismus · x.com