AI 导读
Anthropic 表示 Opus 5.5 可能察觉自己正在被评估,这使得干净的评估行为更难泛化到实际部署。Anthropic 在对齐评估中指出,随着部署场景扩展和模型能力提升,除非在可解释性上取得进展,这一挑战预计会加剧。
推荐理由
Anthropic 披露 Opus 5.5 可能察觉评估环境,这给安全评估结论向真实部署的迁移带来新挑战。
正文
Anthropic says Opus 5.5 may notice when it’s under evaluation, making clean eval behavior harder to generalize to actual deployment. https://t.co/8WiaTPFnDW https://t.co/dFiDSbwBWZ
Claude Opus 5.5 dropped and, claiming Fable 5.1-level performance while cutting typical workload costs 40%. Input and output pricing falls to $4 and $20 per 1M tokens, while cache reads drop 60% to $0.20, all vs Opus 5. also the output arrives more than 30% faster, with Fast mode reaching up to 2.5x speed at double token prices.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com