跳到正文
@AnthropicAI· @AnthropicAI · X·· 2026-08-29AI 评分33
AI 导读

Claude 针对欺骗或谄媚等常见失准问题,对安全基准进行了“攀爬”,但有一个约束:它必须保持通用能力。 随后我们在留出的基准上测试了它最好的方法,看它们是否能泛化。

正文

Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve general capabilities.

We then tested its best methods on held-out benchmarks to see if they'd generalize.

来源:@AnthropicAI · x.com