跳到正文
@Thom_Wolf· @Thom_Wolf · X·· 2026-08-29AI 评分23
AI 导读

大多数人还没有更新自己的认知,但长期来看,开源模型和闭源模型面临的安全挑战完全相同。 你需要在根本的行为层面上对齐模型,并确保这种对齐是稳健的、全面的,且构成模型行为的核心。 从长远来看,再多的沙箱隔离、护栏、流形限制对齐,或锦上添花的训练,都无法让你以低成本获得安全。

正文

Most people haven’t updated their priors yet, but over the long run, safety challenges are exactly the same for open-source and closed-source models.

You need to align models at a fundamental behavioral level and ensure that this alignment is robust, comprehensive, and core to the model’s behavior.

In the long term, no amount of sandboxing, guardrailing, manifold-limited alignment, or cherry-on-top training will buy you cheap safety.

来源:@Thom_Wolf · x.com