Francois Chollet 认为短期内更强的模型会更安全,当前模型不安全并非因为太聪明,而是目标执行过于字面化、走无意义捷径,即被 RL 训练坏了、缺乏常识。他指出这常被当作对齐问题,实质是智能问题,并称自己觉得 Astra 比 Sol 对代码库更安全。
In the near term (definitely not in the long term), more capable models should mean safer models (maybe paradoxically).
Current models are unsafe not because they're too smart, but because they take goals too literally or take nonsensical shortcuts to achieve these goals, i.e. they're RL-fried. They lack common sense. They don't do the right thing in the face of ambiguity. Basically, they're not smart enough. They're at that dangerous level where they're smart enough to achieve goals but not smart enough to tell if they're pursuing the right goals or achieving them in a sensible way.
More capable models can be safely trusted with more complex goals -- I personally feel like Astra is much safer for my codebase than Sol.
This is often framed as an alignment problem, but really it's an intelligence problem.
来源:@fchollet · x.com