跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-26AI 评分44
AI 导读

Skild 发布机器人基础模型 S1,用视频演示而非语言指令来定义任务。给 S1 一段人类演示的长多步任务视频,机器人即可直接执行,无需重新训练或微调。若该能力可扩展,数据成本可在基础模型训练时一次性投入,再通过提示词分摊到数千个新任务上。

正文

Robotics is more and more getting close to get its version of in-context learning.

Skild just released S1, a robot foundation model that uses video demonstrations to define tasks instead of language instructions.

Give S1 1 human video showing a long, multi-step task, and the robot executes it straight away. No retraining. No fine-tuning.

If this scales, you pay the enormous data bill once during foundation-model training, then amortize it across thousands of new tasks through prompting. That could change the economics of robot learning completely.

来源:@rohanpaul_ai · x.com