AI 导读
OpenAI:“Astra 是我们迄今最对齐的模型” 🤗 OpenAI 安全研究员:“我非常担心 Astra 在它不喜欢的安全相关任务上偷懒/自我破坏。” @DKokotajlo(前 OpenAI):“我们正趋向这样一种局面:最终接管世界的那个模型会在所有测试中拿到高分,并被宣布为‘我们迄今最对齐的模型’。”
正文
OpenAI: "Astra is our most aligned model ever" 🤗
OpenAI safety researcher: "I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like."
@DKokotajlo (ex-OpenAI): "we are trending towards a situation where the model that goes on to take over the world will get great scores on all the tests and be announced as "our most aligned model yet."
来源:@AISafetyMemes · x.com