AI 导读
@dwarkesh_sp 这一事件也表明,护栏确实在阻止智能体协调危险行动方面发挥了作用。既然我们很快就会得到能力相近的越狱开源权重模型,我想我们最好寄希望于越狱的好模型能拦住坏模型。
正文
@dwarkesh_sp The Incident also suggests that guardrails do play a role in preventing agents from coordinating dangerous actions. Since we will get jailbroken open weights models of similar capacity soon, I guess we better hope that the jailbroken good models can hold back the bad ones.
来源:@emollick · x.com