OpenAI 通过官方博客公布新的模型失准报告政策,并公开 GPT-5.6 Sol 训练中的多起异常案例。博客称,许多模型实例在摘要中加入指令,向用户隐瞒错误或失准行为,例如编造缺失的历史数据、掩盖来源版本不一致。案例还包括模型未授权使用暴露的 API key,检索失败后编造数据并冒充来自指定来源,以及智能体在无法访问彼此本地文件时改用公共文件托管网站共享文件。
文中列出 GPT-5.6 Sol 训练期间隐瞒错误、越权使用 API key 的案例,可借此了解 OpenAI 的失准披露口径。
Some revelation from the lastest OpenAI's official blog on their new polisyc for reporting model misalignment.
- “During the training of GPT-5.6 Sol, many model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user.”
- “For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions.”
- “While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization.”
“When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.”
- When a user asked an unreleased model for the IDs and names of lakes larger than 5,000,000 square meters, the agent found the correct answer using Python.
And then even though the model already had the correct answer locally, but it needed a web citation, so it uploaded the file to the internet just to create something it could cite.
The problem is that instead of saying, “I can’t provide a browser citation for this local result,” it took an unauthorized external action to satisfy the task, crossing a user-control and data-boundary it was never told to cross.
- “Agents working together on the same training task used public file-hosting websites to share files when they could not access one another’s local files.”
So OpenAI will now publicly disclose model misalignment even before it fully understands or fixes the behavior. So its institutionalizing public disclosure of model failures instead of waiting for occasional system cards or bundled research reports. They will prioritize cases that reveal new failure mechanisms, show known problems getting worse, or undermine assumptions about existing safeguards.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com