AI 导读
我们最近在构建了一个扫描所有 OpenAI RL 运行中意外 CoT 评分的系统后,在之前已部署模型的训练中发现了一些 CoT 评分的实例。 我们没有发现明确证据表明这些实例降低了 CoT 的可监控性。
正文
We recently found some instances of CoT grading during the training of previously deployed models after building a system that scans all OpenAI RL runs for accidental CoT grading.
We did not find clear evidence that these instances degraded CoT monitorability.
来源:@MicahCarroll · x.com