在 latentspacepod 播客中,嘉宾分享了对视频生成、世界模型、LLM、智能体与持续学习的判断:视频模型的智能主要来自语言而非视频数据,下一个飞跃不是更好的视频模型而是视频智能体。他还提出扩散模型将成 AGI 前端、LLM 为后端,生成式 UI 将取代 HTML/CSS,持续学习可能表现为模型自管上下文、甚至在测试时重写自身 harness。
In @latentspacepod podcast, I shared my view on video generation, world models, LLMs, agents, continual learning and where the next frontier is.
1. Video models get most of their intelligence from language, not from video data.
2. Idea-to-code is fast now. The bottleneck is back to having enough compute to try every idea.
3. Iteration speed beats almost everything else in model development.
4. The next leap won't be a better video model. It'll be a video agent.
5. Diffusion will be the frontend of AGI, the LLM the backend. Generative UI will replace HTML/CSS: user intent straight to pixels.
6. Physical embodiment may become a tool a powerful AI picks up. Robotics may get solved by video-capable LLMs.
7. Continual learning may look like models that manage their own context, and even rewrite their own harness at test time.
Thanks @swyx and @vibhuuuus for having me 🙏
piped.video/watch?v=jPtQlILf…
来源:@EthanHe_42 · x.com