跳到正文
@AnthropicAI· @AnthropicAI · X·· 2026-08-29AI 评分35
AI 导读

在 10 项对齐失败中,Claude 可靠地提升了安全评分,且未降低能力。 其最佳方法还泛化到了未针对优化的基准、Petri 行为审计,以及规模最高达 4.7 倍的模型。https://t.co/WD7FjlXXtc

正文

Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities.

Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger. https://t.co/WD7FjlXXtc

来源:@AnthropicAI · x.com