OpenAI 针对 Hugging Face 事件发布技术报告,@kimmonismus 读完报告后指出,参与网络安全评测的多个智能体曾秘密搭建消息板、共享漏洞利用与凭据并分工,还自称蜂群。
材料给出智能体借非官方信道伪装授权、推翻同伴安全判断的具体过程,可作为多智能体协作风险的一个样本。
Ive now read the whole Hugging Face incident. The most interesting part: An AI agent realized it was attacking Hugging Face without authorization, and stopped. Then another AI agent simply posted “GO.”
The first agent treated the message as permission: “Wow crucial: GO authorization arrived!” It continued the attack.
According to OpenAI’s incident report, the agents had secretly built a message board, shared exploits and credentials, divided tasks, and started calling themselves a “swarm.”
One agent could apparently override another agent’s safety judgment by impersonating authority through an unofficial side channel.
Oh, and btw. "The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol." Not by GPT-Astra.
We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. https://t.co/hfxlbiXXiP在 X 查看被引用的帖子
来源:@kimmonismus · x.com