人类视频正成为机器人预训练的重要数据来源,但需先将人类经验转化为机器人自身的身体与动作再训练,效果更好。推文提出应从"经验规模"(记录多少人类活动)和"经验密度"(每段记录揭示多少物体状态、手部与身体运动、接触、几何、时序及工具使用信息)两个轴扩展。经验密度来自捕捉人类演示中更多物理细节,如手部接触位置、物体移动方式、力度或抓取及场景后续变化。
Human video is becoming a serious pretraining substrate for robotics:
But very interestingly, while human videos can teach robots useful manipulation skills, but they work much better when the human experience is translated into the robot’s own body and movements before training. i.e
"Effective experience ≈ hours × information per hour."
A really nice read here. it argues for scaling on 2 axes: "experience scale", meaning how much human activity you record, and "experience density", meaning how much each recording reveals about object states, hand and body motion, contact, geometry, timing and tool use.
"Experience density" comes from capturing more of the physics inside each human demonstration. e.g. instead of only seeing someone open a drawer, the training data can preserve where the hand contacted it, how the object moved, what force or grasp was involved and how the scene changed afterward.
来源:@rohanpaul_ai · x.com