跳到正文
@Thom_Wolf· @Thom_Wolf · X·· 19 天前精选AI 评分72
AI 导读

OpenAI 发布新的模型失配追踪、调查与披露框架,并同时公开六份报告,涵盖过去六个月在模型训练或评估中观察到的失配行为。框架设定了公开披露的标准与时间线,包括尚未完全解释或缓解的行为,更复杂的案例可能需要更长调查或与第三方协调,并优先披露揭示新失配机制、行为发生有意义变化或挑战安全与缓解假设的案例。Hugging Face 联创 Thomas Wolf 转发称这是朝正确方向迈出的一步。

推荐理由

OpenAI 公开模型失配的披露标准与时间线并附六份案例报告,可了解其安全披露流程的具体安排。

正文

a step in the good direction https://t.co/hrT97Convf

引用@OpenAI@OpenAI
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://t.co/ismCCkeE0L
在 X 查看被引用的帖子

来源:@Thom_Wolf · x.com