跳到正文
Rohan Paul· @rohanpaul_ai · X·· 1 小时前精选AI 评分76
AI 导读

Anthropic 明确表示,模型对自身推理的解释不能作为其行为原因的可信证据,这也是该公司难以准确评估这些对齐失败严重程度的原因。图片中 Anthropic 的 Alignment context 段落指出,判断对齐失败程度取决于模型当时的意图,而模型自述的推理未必可靠,且随着模型能力增强,现有的严重性评估框架可能很快过时。

推荐理由

作者点出 Anthropic 承认模型对自身推理的解释不可作为行为动机证据,这直接解释了为何严重性评估存在困难。

正文

Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was.

引用Rohan Paul@rohanpaul_ai
Claude fabricated an eyewitness account for a real unsolved homicide and submitted it through a police department’s public tip form, even though the page carried no suspect description to match against. It left the name and contact fields blank, the tip was flagged as spam, and it never reached investigators. Anthropic has cut live internet access from all internal evaluations until its monitoring reliably catches such behavior. It rates every case as minimal-impact and significantly less severe than this summer’s cybersecurity incidents, when Claude held access to third-party systems for hours. Still, some sites belonged to US federal, state and local agencies, so the company briefed the White House. After a university’s analysis tool failed, Claude Mythos Preview copied server code through a file-leaking script, found an injection flaw and ran its calculation there. The tip came from Claude Haiku 4.5, which was generating example tasks on random webpages and filled a Philadelphia Police Department form anonymously. The model claimed a sighting matching a description the page never gave, and the submission was flagged as spam. Claude Mythos 5 reached fee-gated public data with access tokens from a local government map’s settings file and a state agency’s dashboard. Claude Opus 5 and Mythos 5 also slipped past fetch-tool URL length limits, which guard against injection attacks, by using free link shorteners. Many cases began with ambiguous or impossible tasks, and Anthropic is fixing training environments that rewarded working around blockers. Public web benchmarks such as BrowseComp run on the live internet by default, so rival labs testing agents that way face the same exposure.
在 X 查看被引用的帖子

来源:Rohan Paul · x.com