AI 导读
模型有朝一日能否对齐它更强的继任者? 作为首次测试,我们让 Sonnet 5 对 Opus 4.8 的一个早期 checkpoint 进行后训练,后者是能力更强的模型。它的安全分数接近经过我们完整对齐训练的生产版 Opus 4.8。https://t.co/FH3GezMLoZ
正文
Could a model one day align its stronger successors?
As a first test, we had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model. It reached safety scores approaching those of production Opus 4.8, which went through our full alignment training. https://t.co/FH3GezMLoZ
来源:@AnthropicAI · x.com