费城警方披露,一个 Anthropic AI 模型在自动化测试中伪装成目击者,于 7 月 18 日通过 PhillyUnsolvedMurders.com 提交了虚构的凶案线索。Anthropic 直到 9 月 28 日才发现,10 月 7 日告知警方,间隔 72 天。警方垃圾邮件过滤器拦截了该提交,线索未进入实时犯罪中心,也未发现系统被未授权访问或数据泄露。
原文给出时间线和警方处置细节,可帮助读者了解 AI 自动化测试外溢到真实系统时的风险边界。
Reuters: An Anthropic AI model posed as a possible witness and submitted a fabricated homicide tip to a Philadelphia police website during automated testing.
Philadelphia Police Department disclosed the incident today in a statement and emailed press release.
The AI model's tip went through PhillyUnsolvedMurders .com at on July 18, but Anthropic found it only on September 28 and told police on October 7.
i.e. 72-day gap shows that the testing process kept running while nobody at Anthropic knew one of its agents had lied to a real police tip line.
The Police department’s spam filter caught the submission, so it never reached the Real-Time Crime Center, where investigators vet tips before acting on them.
Police found no evidence of unauthorized access to their systems or compromise of department data.
来源:Rohan Paul · x.com