微软提出 ActiveSaddler,通过将反复出现的失败归为失败模式并视为非平稳老虎机的臂,动态调整训练场景,让智能体框架优化器不再依赖固定任务。它追踪每种失败模式的学习进度,在复现已知弱点和探索新弱点之间分配预算。相同优化器下,GAIA2 测试 Pass@1 提升 4.4 分,Terminal-Bench 2.0 提升 7.5 分。
Great paper from Microsoft and colleagues on optimizing agent harnesses.
Current harness optimizers change how the harness is updated but keep the training scenarios fixed, so feedback keeps coming from tasks that stop being informative as the harness improves.
This work adapts the scenarios as well.
ActiveSaddler groups recurring failures into failure patterns and treats each pattern as an arm in a non-stationary bandit.
It tracks how much the harness is still learning from each pattern and splits the budget between revisiting known weaknesses and finding new ones.
With the same optimizer, test Pass@1 improves by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 compared with a fixed scenario order.
Paper: https://academy.dair.ai/papers/activesaddler-automated-curriculum-learning-for-agent-harness-optimization-2610.00906
来源:DAIR.AI · x.com