Luma AI 联合创始人 @gravicle 在访谈中提出,过去 5 年由 LLM 主导,未来 5 年将是视觉智能的时代。他回顾了从 South Park Commons 到 Luma Labs 再到 world model 的历程,并讨论了 Luma Agents 与端到端创作流程、统一模型为何优于纯视频 scaling,以及消费级视频生成为何失败。
The last 5 years have been dominated by LLMs.
@gravicle makes a compelling case for visual intelligence dominating the next 5.
We discussed the journey from @southpkcommons, to @LumaLabsAI, to world models (what does this even mean!?).
Enjoy!
(00:00) Apple's LiDAR work seeded Luma's founding vision
(05:05) - A free app was secretly a data collection operation
(06:49) H100s made video generation stop feeling impossible
(10:00) Most definitions of "world model" are simply wrong
(14:37) Why unified models beat pure video scaling
(20:03) Luma Agents and the end-to-end creative production loop
(23:00) The knowledge gap no model can close alone
(29:34) Blowing James Cameron's mind in five minutes
(31:16) Why consumer video generation failed—and who it's actually for
(35:23) - Product and research aren't independent things at a foundation lab
Video
来源:@finn_meeks · x.com