跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 14 天前AI 评分58
AI 导读

Rohan Paul 指出,根据 Claude Opus 5.5 系统卡,METR 另有一个对 Anthropic 内部 AI 研发数据拥有更深访问权限的团队,该团队只向公开评估团队分享了结论,没有分享支撑证据或背后的推理过程。因此,对 Anthropic 由 AI 驱动研发加速的公开评估,有一部分建立在外部人乃至另一个 METR 团队都无法独立核查的证据之上。

正文

METR had a separate team with deeper access to Anthropic’s internal AI R&D data. That team shared its conclusions with the public-assessment team, but not the supporting evidence or reasoning behind them.

- from Claude Opus 5.5 system card.

i.e. part of the public assessment of Anthropic's AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.

引用@rohanpaul_ai@rohanpaul_ai
Some revelation from the Claude Opus 5.5 system card. - Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text - Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place. - Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year. - Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real. - Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs. "During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs" - METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com