一项研究分析大量公开后训练轨迹后发现,智能体的训练策略在第一步就被锁定,剩余全部预算都花在该策略内的局部调整上。三种递进修复中,经验驱动脚手架在 GSM8K 上提升 12.6 分、HumanEval 上提升 40.8 分,但策略依然冻结;人类引导能改变开局选择,训练开始后智能体仍滑回局部循环;额外推理算力只在简单任务上有效。智能体缺少的是在执行过程中重新考虑策略的能力。
Great paper if you are tracking progress in recursive self-improvement (RSI).
(bookmark it)
There is so much hype around RSI, so I think it's worth understanding why current models are not able to do this properly yet.
Issues range from "lack of creativity" of models to getting stuck in a local optimum.
This work tries to provide more insights into whether agents can really post-train other agents.
Here is the most interesting finding reported in the paper: "the agent’s training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy."
They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it.
They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen.
Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.
What agents lack here is a way to reconsider strategy while execution is still running.
Paper: https://t.co/3XJAfRYmtB
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com