跳到正文
@latentspacepod· @latentspacepod · X·· 2026-06-01AI 评分45
AI 导读

前 xAI 世界模型负责人、前 NVIDIA Cosmos 研究员 Ethan He 在 latent.space 播客中提出,AI 视频可能重走编程智能体的路径,文本生成视频只是"自动补全"阶段。他复盘了 Grok Imagine 从 0 到 1 的过程,并认为世界模型将走向实时交互,语言模型或成为视频的控制层,未来 AI 视频更像一个带摄像头、剪辑器、时间线和工具带的智能体,而非提示词输入框。

正文

🆕Grok Imagine’s Video Agent Moment: Cosmos, xAI, World Models, Generative UI, & the Codex Phase for Video!

latent.space/p/video-agents

@EthanHe_42, former @xai world model lead and @nvidia Cosmos researcher, explains why AI video may follow the same path as coding agents, how Grok Imagine went from zero to one, why text-to-video is only the autocomplete phase, how world models become real-time and interactive, why language models may become the control layer for video, and why the future of AI video may look less like a prompt box and more like an agent with a camera, editor, timeline, and tool belt.

来源:@latentspacepod · x.com