一篇论文分析大量公开后训练轨迹后发现,智能体在第一步就锁定训练策略,剩余预算全用于局部调整。经验驱动的脚手架将 GSM8K 提升 12.6 分、HumanEval 提升 40.8 分,但策略仍被冻结;人类引导可改变开局选择,训练开始后智能体又滑回局部循环。增加推理算力只在简单任务上有效,对最难任务几乎无用。
Finally a good paper testing whether agents can really post-train other agents.
(bookmark it)
They analyzed a large corpus of publicly released post-training trajectories. Across tasks, the agent locks in its training strategy at the very first step and spends the entire remaining budget on local adjustments inside it.
They then tried three escalating fixes. An experience-driven scaffold lifted execution broadly, worth 12.6 points on GSM8K and 40.8 on HumanEval, and the strategy stayed frozen.
Human guidance redirected the opening choice, and the agent slid back into local loops once training began. Extra inference compute paid off on easy tasks and did almost nothing on the hardest one.
What agents lack here is a way to reconsider strategy while execution is still running.
Paper: https://t.co/3XJAfRYmtB
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com