AI 导读
OpenAI 公布了一套追踪、调查并披露模型失配行为的框架,并同时发布过去六个月内在其模型训练或评测中观察到的六份失配报告。框架设定了公开披露的标准和时间线,包括尚未完全解释或缓解的行为,更复杂的案例可能需要更长时间调查或与第三方协调。披露将优先选择揭示新失配机制、已知行为出现实质性变化,或挑战现有安全与缓解假设的案例。OpenAI 表示这是起点,后续会依据经验和公众反馈完善流程并持续发布报告。
推荐理由
OpenAI 把模型失配的披露从偶发系统卡变为常态化流程,读者可了解其公开标准与优先顺序。
正文
So OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior.
So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports.
They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://t.co/ismCCkeE0L在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com