跳到正文
@OpenAI· @OpenAI · X·· 20 天前精选AI 评分67
AI 导读

OpenAI 发布一套用于追踪、调查和披露模型失准(misalignment)实例的框架,并同步公开近六个月在其模型训练或评估中观察到的六份失准行为报告。框架设定了公开披露的标准和时间线,包括尚未完全解释或缓解的行为,更复杂的案例可能需要更长时间调查或与第三方协调。

推荐理由

OpenAI 给出了模型失准行为的披露标准、时间线和六份案例,可供观察其安全流程的透明度。

正文

We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.

The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.

We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.

Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.

This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.

https://t.co/ismCCkeE0L

来源:@OpenAI · x.com