AI 导读
我们可能只剩几个月了。 OpenAI 顶尖研究员之一 Noam Brown 刚刚表示,再发布 1 到 2 个模型之后,模型的研究品味可能就会超过他。 换句话说,RSI。
正文
We may have just a few months left.
Noam Brown, one of OpenAI's top researchers, just said models 1 or 2 releases from now may have better research taste than him
In other words, RSI. https://t.co/BcuGjewfG2
OpenAI researcher says slowing down is not enough: "a ticking time bomb" "Models will increasingly seem aligned even when they are not. The models will likely convince people that everything is fine." "The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power."在 X 查看被引用的帖子
来源:@AISafetyMemes · x.com