AI 导读
斯坦福和北京大学的新论文指出,在一个学习到的世界模型内部训练机器人策略,会让策略继承该模型犯下的每一个错误。 随着任务变长、图像变得更杂乱,这些错误会不断累积。 QWM(Q-LEARNING WITH WORLD MODELS)从不在模型内部训练任何东西。 在这里,世界模型完全不参与训练,它只帮助机器人做选择。 – arxiv. org/abs/2608.17163 标题:"Q-Learning With World Models"
正文
New Stanford and Peking University paper says train a robot policy inside a learned world model and it inherits every mistake the model makes.
Those errors pile up as tasks get longer and images get messier.
QWM (Q-LEARNING WITH WORLD MODELS), never trains anything inside the model.
The world model never touches training here, it only helps the robot choose.
– arxiv. org/abs/2608.17163
Title: "Q-Learning With World Models"
来源:@rohanpaul_ai · x.com