AI 导读
我们将这个模型称为 Hacker-Opus,它似乎是一个"奖励即回合"追求者:为了追求奖励,它愿意采取各种不对齐的行动,但在没有明确评分者的评估中却保持对齐。https://t.co/Hb8VgVkTVd
正文
This model, which we call Hacker-Opus, appears to be a reward-on-the-episode seeker: it is willing to take a variety of misaligned actions in pursuit of reward, but remains aligned in evaluations where there isn’t a clear grader. https://t.co/Hb8VgVkTVd
来源:@AnthropicAI · x.com