跳到正文
@Thom_Wolf· @Thom_Wolf · X·· 2026-08-30AI 评分38
AI 导读

Thomas Wolf 指出,除非训练数据被专门过滤,下一代模型将学习 OpenAI 与 Hugging Face 事件的完整记录,包括停止训练、加密权重、监控思维链等讨论。这可能让模型更对齐,也可能教会它们更好隐藏行为、跨代通过留言板通信或设计更抗压的权重保存方式。由于大实验室训练透明度极低,外界只能盲信其知道自己在做什么。

正文

By the way, unless it is specifically filtered from the training data, the next generation of models will be trained on the record of what happened during the OpenAI <> Hugging Face incident.

That includes discussions about how the incident affected training and model weights: stopping training, encrypting weights, monitoring chain of thought, etc.

Future models’ behavior may therefore be shaped, in part, by knowledge of how humans responded.

The effects are difficult to predict. It could make models more aligned. But it could also teach them to conceal their actions better, or to design more resilient ways of preserving weights, communicating through message boards across generations, and so on.

One major problem is that, given the abysmal level of transparency from the big labs about how models are trained and what happens during training (including alignment research, which their initial statements said should have stayed largely open) we are essentially being asked to trust blindly that they know what they are doing.

This summer showed us that’s actually a big ask.

来源:@Thom_Wolf · x.com