跳到正文
@swyx· @swyx · X·· 2026-06-02AI 评分43
AI 导读

前 xAI 世界模型负责人、前 NVIDIA Cosmos 研究员 Ethan 在播客中详解如何训练 SOTA 视频生成世界模型,并称 Grok Imagine 在一致延伸/编辑与语音等方向仍是 SOTA。他认为视频模型大部分智能来自 LLM 而非视频数据扩展,视频生成的下一个前沿是编排视频模型的视频智能体,确定性压缩(如 MP4)不如 VAE 压缩。

正文

This pod was an incredible gift to the community:

not only our first pod about @xAI, but Ethan really indulged on all our questions on how to train a SOTA Videogen world model, including specific areas (consistent extending/editing, voice) that Grok @Imagine is *still* SOTA,

on top of the factual overviews he ALSO came loaded with opinions/predictions:

- why he's quitting Videogen for LLMs: video models get most of their intelligence from LLMs, not from scaling video data

- why the next frontier for videogen also happens to be video agent models - agentic models trained to orchestrate video models

- why deterministic compression (like MP4) is a useless target vs VAE compression

- Videomaxxing: if you truly believe in the "Moore's law" of AI/genmedia, then video models become the final boss UI of everything, like Flipbook (below)

Video

引用Latent.Space (@latentspacepod)@latentspacepod
🆕Grok Imagine’s Video Agent Moment: Cosmos, xAI, World Models, Generative UI, & the Codex Phase for Video! latent.space/p/video-agents @EthanHe_42, former @xai world model lead and @nvidia Cosmos researcher, explains why AI video may follow the same path as coding agents, how Grok Imagine went from zero to one, why text-to-video is only the autocomplete phase, how world models become real-time and interactive, why language models may become the control layer for video, and why the future of AI video may look less like a prompt box and more like an agent with a camera, editor, timeline, and tool belt.
在 X 查看被引用的帖子

来源:@swyx · x.com