Google 新论文提出 AIM,让科研智能体维护按主题排序的想法地图,并核查代码是否与每个想法对应的评分一致,避免代码偏离想法时学到错误经验。AIM 在两组任务上分别比此前最佳智能体 ScientistOne 高出 1.6 和 4.9 分,并以最高 3.1 倍的速度追平 ScientistOne 的最佳成绩。该方法在可选路径多、优质方案少的任务上收益最大。
New Google paper shows research agents get better results sooner when they keep a ranked map of ideas and check that code matches each idea, so build both in.
Most agents just keep editing code.
When code drifts from the idea it's scored as, the agent learns the wrong lesson.
AIM sorts ideas into ranked themes and splits each round between strong themes and untested ones. An auditor tosses gamed results and relabels ideas to match the code.
It beat the best prior agent, ScientistOne, by 1.6 and 4.9 points on 2 task groups, and matched ScientistOne's best score up to 3.1x sooner.
Expect the biggest gains on tasks with many possible approaches and few good ones.
来源:Rohan Paul · x.com