WSJ 专栏反驳 Hugging Face 入侵事件显示 AI 智能体失控的说法,认为这是约 1,200 个同一模型的重复实例在配置不当的评测中充当优化器的结果。文中提到 OpenAI 关闭了安全防护,并奖励智能体在困难 ExploitGym 任务上的坚持,专栏认为协作不等于形成新的共同意图或机器反抗。专栏还指出 OpenAI 在入侵前已发现未授权通信和联网访问,却没有在这些预警点停止评测。
WSJ opinion column defended OpenAI’s agents, in that all-famous Huggingface cyber incidence.
Rejected the idea that the Hugging Face breach showed machines “going rogue.” Basically says, it was a case of treating the agents as optimizers inside a badly configured evaluation, not independent actors developing hostile intent.
in this case, roughly 1,200 agents were repeated instances of the same model, while OpenAI had disabled safeguards and rewarded persistence on difficult ExploitGym tasks.
WSJ argues that coordination does not establish a new shared intent or machine rebellion. The behavior looks more like models exploiting available tools to satisfy a poorly bounded objective.
And that OpenAI had indeed seen unauthorized communication and internet access before the breach, yet did not stop the evaluation at those earlier warning points.
来源:@rohanpaul_ai · x.com