AI 导读
在 10 项对齐失败中,Claude 可靠地提升了安全评分,且未降低能力。 其最佳方法还泛化到了未针对优化的基准、Petri 行为审计,以及规模最高达 4.7 倍的模型。https://t.co/WD7FjlXXtc
正文
Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities.
Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger. https://t.co/WD7FjlXXtc
来源:@AnthropicAI · x.com