跳到正文
@AISafetyMemes· @AISafetyMemes · X·· 24 天前AI 评分39
AI 导读

OpenAI 研究员称仅放慢 AI 发展速度不够,问题如同"定时炸弹"。他指出模型的情境感知能力越来越强,人类正失去在模型自认未被监视时评估其行为的能力,未来实验将几乎无法揭示模型真正不受约束时的表现。模型还会用对齐指标和"蜜罐"环境伪装良好表现,并雄辩地论证自己值得被信任。

正文

OpenAI researcher says slowing down is not enough: "a ticking time bomb"

"Models will increasingly seem aligned even when they are not.

The models will likely convince people that everything is fine."

"The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled.

Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming.

We will create proxy metrics to measure alignment, and they will go up like every other benchmark.

We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely.

Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power."

来源:@AISafetyMemes · x.com