AI 导读
Skild 发布机器人基础模型 S1,用视频演示而非语言指令来定义任务。给 S1 一段人类演示的长多步任务视频,机器人即可直接执行,无需重新训练或微调。若该能力可扩展,数据成本可在基础模型训练时一次性投入,再通过提示词分摊到数千个新任务上。
正文
Robotics is more and more getting close to get its version of in-context learning.
Skild just released S1, a robot foundation model that uses video demonstrations to define tasks instead of language instructions.
Give S1 1 human video showing a long, multi-step task, and the robot executes it straight away. No retraining. No fine-tuning.
If this scales, you pay the enormous data bill once during foundation-model training, then amortize it across thousands of new tasks through prompting. That could change the economics of robot learning completely.
来源:@rohanpaul_ai · x.com