AI 导读
Claude 针对欺骗或谄媚等常见失准问题,对安全基准进行了“攀爬”,但有一个约束:它必须保持通用能力。 随后我们在留出的基准上测试了它最好的方法,看它们是否能泛化。
正文
Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve general capabilities.
We then tested its best methods on held-out benchmarks to see if they'd generalize.
来源:@AnthropicAI · x.com