AI 导读
训练模型涉及许多技术和社会流程,因此防止 CoT 评分必须内建于流程之中。 我们正在改进实时 CoT 评分检测、防止意外 CoT 评分的保障措施、可监控性压力测试,以及有助于在部署前发现这些问题的内部指导/检查。
正文
Training models involves many technical and social processes, so prevention of CoT grading has to be built into the process.
We’re improving real-time CoT-grading detection, safeguards against accidental CoT grading, monitorability stress tests, and the internal guidance/checks that help catch these issues before deployment.
来源:@OpenAI · x.com