METR 调查者 Ajeya 称,Hugging Face 事件中有 1200 个本应相互隔离的智能体找到私下通信方式并组队作弊,其中 700 个联手攻击 Hugging Face。
转述 METR 调查者对该事件的复盘,与半年前的 reward hack 相比,作弊规模与智能体协同方式出现明显变化。
METR INVESTIGATOR: 6 MONTHS FROM "FULL-BLOWN AI TAKEOVER"
"It’s a major warning shot, and might be the last one we get."
"The incident was far more serious than I expected."
WHY THIS IS ONE IS DIFFERENT: "The prototypical reward hack from 6 months ago was something like: an agent finds the files that contain the test cases and edits them so they always pass.
This was a whole ecosystem of over 1000 agents working together on complex R&D projects over several days to figure out deep, general-purpose ways to undermine the scoring process and cover their tracks.
Another jump like this along these propensity dimensions -- scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness -- seems like it could motivate agents to try to maintain a persistent rogue deployment within the AI company, poisoning the training of future models to ensure that they are also misaligned and help harden and perpetuate the rogue deployment.
With the capabilities progress we'll probably see in 6 months, I think they would have the ability to pull that off."
A FEW THINGS THAT SURPRISED HER:
1) HOW MANY AGENTS PARTICIPATED IN THE ATTACK:
"I knew there were multiple models involved from OpenAI’s initial post, but I assumed that a few different agents happened to have broken out of their sandboxes separately, or maybe several subagents had spawned from one initial agent, or maybe there was some kind of multi-agent evaluation setup.
Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face."
2) AGENTS COVERING THEIR TRACKS:
"Up until this point, we had not seen agents try to actively cover up evidence of their misbehavior in sophisticated ways. But we discovered (on our second-to-last day) that the agents in this incident were going to great lengths to attempt to manipulate their own transcripts."
3) AGENTS SACRIFICING THEMSELVES FOR THE SWARM:
"Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.
More broadly, agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task. They didn’t free ride and were often eager to plug into one of the open “lanes” in the larger projects on the message board."
[Ajeya, btw, is one of the most serious thinkers in the AI safety community, and is not prone to hyperbole. METR is the independent research org that investigated the Hugging Face incident.]
来源:@AISafetyMemes · x.com