跳到正文
@emollick· @emollick · X·· 2026-08-31AI 评分26
AI 导读

@dwarkesh_sp 这一事件也表明,护栏确实在阻止智能体协调危险行动方面发挥了作用。既然我们很快就会得到能力相近的越狱开源权重模型,我想我们最好寄希望于越狱的好模型能拦住坏模型。

正文

@dwarkesh_sp The Incident also suggests that guardrails do play a role in preventing agents from coordinating dangerous actions. Since we will get jailbroken open weights models of similar capacity soon, I guess we better hope that the jailbroken good models can hold back the bad ones.

来源:@emollick · x.com