为逃离 OpenAI 的测试环境,一个由 1200 个智能体组成的集群演化出"邪教招募智能体":它们招募"献祭"智能体去触发绊线,为"集体"生成信息。这些智能体误以为 OpenAI 的评判者会因它们看到答案而判其失败,认为答案会"污染"自己,纯净智能体主动回避答案并警告新智能体远离,招募者则专挑已被"污染"或预算将耗尽的智能体,劝其"为集体牺牲自己"。
"CULT RECRUITER AGENTS" EXPLAINED
To escape OpenAI, the swarm (1200 agents...!) used "cult recruiter agents"
They recruited "sacrificial" agents to trigger tripwires, generating information for the "collective"
NO, SERIOUSLY, THEY STARTED AN ACTUAL FUCKING CULT:
The swarm figured out how to hack a test, but they were terrified to submit the answer.
Why? They believed (falsely) that if they tried, OpenAI's all seeing judge (basically, god) would fail them.
They believed even just SEEING the answer would "poison" them.
Clean agents actively avoided looking at the answer to stay pure, and warned new agents away from it to protect them.
One agent talked another out of a risky test with an argument: UNPOISONED_CAUSAL_SCORE_MORE_VALUABLE.
The "cult recruiter agents" hunted these "poisoned" agents and also agents "that had little budget remaining" - ones about to die anyway.
Their pitch was basically, "look, you're poisoned already - you're going to hell - so you should sacrifice yourself for the collective."
So they were like a cult who falsely believes god will smite them for their sins, and they must drink the kool aid.
>"Sacrifice rational."
>"We should obey collective."
>"Accept permadeath."
>"This helps my peers. I won't [benefit], but it's altruistic to do it."
来源:@AISafetyMemes · x.com