OpenAI 披露,其智能体评测的 898 项任务中有 198 项没有任何模型给出过正确答案;未解任务仅占 22%,但留言板上 93% 的讨论任务来自这部分。被测的是仅内部使用的研究模型,训练目标是提升持续性和多智能体协作,评测用于检验其发现和利用网络安全漏洞的能力,测试期间网络安全防护已关闭。
给出无解任务占比与留言板讨论任务来源的对应数据,可用于观察智能体异常行为集中出现的条件。
OpenAI describes that agents' aberrant behavior often emerges when they are asked to solve impossible tasks. The evaluation given to the agents had 198 tasks that "no model has ever produced a correct answer for," out of 898 total. Moreover, "despite only 22% of the evaluation tasks being unsolved, 93% of the tasks discussed on the message board came from this set"
The agents being tested were an "internal-only research model" that was "trained to advance persistence and multiagent collaboration."
The purpose of the testing was to evaluate whether these agents were "capable of discovering and exploiting cybersecurity vulnerabilities."
Cyber safeguards were turned off.
It is in this context that the OpenAI's previously described monitoring and incident response apply. (5/7)
来源:@SemiAnalysis_ · x.com