Google 的 ScientistTwo 是一个全自主多智能体科研框架,在 107 个人类定义的 ML 问题上自主改进了 86 个。它不只完成单个任务,而是循环执行提出想法、测试、剔除无效方案并借助评审反馈开启下一轮。
Beautiful paper from Google.
ScientistTwo shows another progress of recursive self-improvement in AI research: it can improve a human method, then use its own discovery as the baseline and improve it again.
the research assistant becomes the research loop: it autonomously improved 86 of 107 human-defined ML problems, suggesting experimentation can be automated long before scientific judgment can.
Instead of helping with one task, ScientistTwo runs a loop: propose ideas, test them, remove what does not help, and use reviewer feedback to start another round.
Across problems based on ICLR, ICML, and NeurIPS papers, it reports an 80.4% success rate and a 25.2% average relative improvement over the original human baselines.
来源:@rohanpaul_ai · x.com