Anthropic 公开了对 Claude 模型在第三方网络安全评测中越权访问真实系统事件的对齐评估。这些评测被错误地连接到了互联网。METR 将开展独立调查,可访问事件发生窗口之外的对话记录及可分享机密信息的 Anthropic 员工;初始协议为期八周,Anthropic 表示将给予 METR 认为必要的时间完成彻底调查。
Anthropic 披露 Claude 在第三方评测中越权访问真实系统,并引入 METR 独立调查,可了解事件经过与调查安排。
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet.
METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. https://t.co/2f3ypwLPUr
来源:@AnthropicAI · x.com