一篇立场论文主张,科学智能体不应只评估智能体本身,而应把"科学家+智能体"作为整体来研究。论文指出当前系统多把人类当作监督者,而 GPT-5-mini 在 10 项科学任务中几乎从不主动求助;案例显示专家能发现智能体遗漏的错误,智能体则加速执行,收益来自协作而非自主。论文据此提出新基准:人机团队能否比任何一方单独产出更好的科学。
The race to build “AI Scientists” may be optimizing for the wrong unit: the agent alone.
This position paper argues that scientific agents should be studied as human-agent systems, where the thing you evaluate is the scientist + agent pair.
Most current systems still treat the human as a supervisor: set the goal, review a phase, approve the final artifact. Far fewer are built for continuous, fine-grained collaboration during the work itself.
The problem is that agents do not naturally know when they need human input. In 10 science tasks, GPT-5-mini almost never asked for help.
But that input matters: in the case studies, experts caught errors the agents missed, while the agents sped up execution. The gain came from collaboration, not autonomy.
The paper’s proposed benchmark is therefore different: does the human-agent team produce better science than either member alone, without collaboration cost overwhelming the gain?
– arxiv. org/abs/2608.14667
Title: "Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems"
来源:@rohanpaul_ai · x.com