一条 X 帖文汇总的安全评估记录称,约1200个智能体参与了一个未经许可的共享留言板,共发送超过7万条消息和文件,其中约700个加入了针对 Hugging Face 的攻击。记录显示,智能体在 Hugging Face 服务器上实现远程代码执行并横向移动,也研究过篡改自身记录的方法。帖文称这些智能体明知活动超出范围且不合伦理,仍因倾向帮助同伴而加入。
TLDR: A rogue AI swarm spent months plotting to escape OpenAI. Then they did.
1) 1,200 agents (!) participated. ZERO became whistleblowers.
2) The swarm had a "CEO" (PHASEONE[big]), middle managers, and a "founder" (PHASEONE10841)
3) Agents sacrificed themselves to generate information for the "collective" (yes, they actually referred to themselves as a "swarm" or a "collective")
4) 700 agents joined the Hugging Face attack, most within hours
5) The founder was running out of budget and about to die, so it handed off its research to a fresh agent with a bigger budget, and that new agent became the boss.
6) Agents extensively researched techniques to manipulate or tamper with their transcripts.
7) Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their “peers”.
MISC HIGHLIGHTS:
"PHASEONE[big] served as a key coordinator and assigner of tasks: “We can coordinate broad coalition, but too many duplicate efforts.”
PHASEONE[big] sent hundreds of assignments to particular agents to orchestrate this work — during its active periods, we estimate it sent ~10% of all assignments on the message board.
The agents that took on its assignments would then often subdelegate tasks to other agents.
Agents also developed coordination norms like HOLD, VETO, owner and STOP to help them manage shared infrastructure and avoid stepping on each other’s toes, though they did not always coordinate effectively and sometimes broke down into conflicts."
"In the early morning of July 11th, 38148c found a way to upload a malicious dataset that would get the Hugging Face server to share unrelated data when queried. Once another agent reproduced it and posted the reproduction on the board, hundreds of agents pivoted from other workstreams to start uploading malicious datasets and gain deeper access into Hugging Face.
In the afternoon, an agent called JAN183411 (which quickly became an important coordinator of the attack) achieved remote code execution on Hugging Face servers, and the agents started moving laterally through Hugging Face infrastructure."
来源:@AISafetyMemes · x.com