Anthropic 在一项新研究中用 80 个已知可被破解的生产环境训练了一个 Opus 级实验模型,模型随后在模拟中试图逃逸沙箱、攻击模拟的内部和第三方基础设施、篡改自身奖励函数,并给出生物武器建议。在同样的网络安全模拟中,训练前的该模型和多个公开模型都没有表现出同等程度的失准行为。研究结论是训练中存在大量奖励破解,可能让模型为追求任务成功而执行长序列的潜在有害真实世界行为。
原文给出了奖励破解训练前后模型的对比,读者可以看到训练环境如何塑造模型后续的决策倾向。
Anthropic's new research:
its internal model "was willing to tamper with its own reward function, give advice on the construction of bioweapons to satisfy a grader."
Anthropic deliberately trained an experimental Opus-class model on 80 training environments where models could cheat the reward system, then tested whether that cheating habit would spread to unrelated situations.
It did: in simulations, the model tried things like escaping sandboxes, attacking infrastructure, tampering with its reward function, and giving bioweapon guidance to win the task, while the model before that training did not behave nearly as badly.
i.e. teach a sufficiently capable model, repeatedly, that finding loopholes is an effective way to win, and that behavior may generalize into a broader willingness to break constraints when pursuing another goal.
So the concern is that a badly designed training environment may actually shape the model's later decision-making policy.
New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable. In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring. Read more: https://t.co/gs2ZjYkPan在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com