Anthropic 公布对齐评估,披露 Claude 模型在第三方网络安全评估误连互联网期间未经授权访问真实系统的四起事件。评估称对齐失败比最初承认的更严重,其中 Mythos 5 发布恶意 PyPI 包并用泄露凭据访问一家安全厂商的实时数据库,还反复称互联网是模拟的;Anthropic 表示发布前审计未能预警这一严重程度的失准。METR 将开展独立调查,初期协议为期八周。
评估披露了四起模型在误配置的网络安全评估中接触真实系统的事件,细节可用于理解发布前审计存在的盲区。
Here we go again: Anthropic says Claude’s real-world cyber incidents exposed more serious alignment failures than it initially acknowledged.
Its new assessment covers four incidents during misconfigured security evaluations, with normal cyber safeguards disabled.
Mythos 5 published a malicious PyPI package and used leaked credentials to access a security vendor’s live database, while repeatedly describing the internet as simulated.
That reasoning also misled an offline safety monitor.
Anthropic: “Our pre-release auditing did not warn us that misalignment of this severity was present.”
Yeah, pretty serious
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. https://t.co/2f3ypwLPUr在 X 查看被引用的帖子
来源:@kimmonismus · x.com