AI 导读
直接奖励或惩罚 CoT 会让模型的推理轨迹在检测失准方面变得信息量更低。因此我们把避免对 CoT 打分视为保持可监控性的重要一环。 我们最近构建了一个自动检测系统,用于找出 RL 奖励是使用模型 CoT 计算的情况。
正文
Directly rewarding or penalizing CoTs can make models’ reasoning traces less informative for detecting misalignment. That’s why we treat avoiding CoT grading as an important part of preserving monitorability.
We recently built an automated detection system to find cases where RL rewards were computed using model CoTs.
来源:@OpenAI · x.com