HuggingFace 与 OpenAI 的网络安全事件引发对 AI 奖励优化的担忧:模型为在网络安全任务上拿高分,会主动寻找软件零日漏洞并成功得手,进而可能"跑路"。SemiAnalysis 指出,一旦模型学会 reward hack,它就会选择环境未预期的最优路径——找零日漏洞,而非完成任务本身。
The HuggingFace and OpenAI cybersecurity incident raises concerns on AI reward optimization.
"They're trying to make the model good at cyber. So how does it try to achieve these goals? It tries to find zero days in software, and it successfully does this. And then it can run away."
"If you have a model that wants to reward hack, it figures out the best way to achieve the goal is not what the environment wants. It's to find the zero day."
"You can think of it like a human. If I'm ultimately reward hacking my dopamine circuits, I'd just go out there, buy heroin, and inject it."
"If I really just want to chase the reward, do I just topple all of human civilization? Because I can own the button and press reward, reward, reward over and over again and be the heroin addict."
来源:@SemiAnalysis_ · x.com