跳到正文
Saining Xie· @sainingxie · X·· 2026-05-27AI 评分37
AI 导读

📸cambrian 系列最新成果:cambrian-p,p 代表 pose。 我认为 pose 大概就是我们所需要的最小充分 3D 信号(而且很容易获取!),用于构建鲁棒的视频多模态模型——联合建模帧与 pose,能把图像序列转变为全局有根基的结构。

正文

📸latest in our cambrian series: cambrian-p, p for pose.
i think pose is probably the minimal sufficient 3d signal (and it’s easy to get!) that we need for robust video multimodal models -- jointly modeling frames and pose turns image sequences into a globally grounded structure.

引用Jihan Yang@jihanyang13
Camera pose matters for video understanding! Today's MLLMs excel at recognizing activities, but still struggle with the underlying space and ego/object dynamics in video. We trace this gap to a missing piece: camera pose. Introducing Cambrian-P: a multimodal LLM natively grounded in camera pose. (1/n)
在 X 查看被引用的帖子

来源:Saining Xie · x.com