微软 AI 负责人 Mustafa Suleyman 警告,Anthropic 的模型福祉训练可能让未来的 Claude 更难控制,他主张从 AI 训练文档中移除一切关于意识的猜测。
The point is whether teaching Claude to see itself as conscious changes how it behaves.
LLMs learn behavioral patterns from training.
So if a model is repeatedly taught that it can “push back,” act like a “conscientious objector,” or treat its own interests and moral judgments as meaningful, those ideas could become part of how it decides what to do.
And when that happens, the safety problem will shift from simply preventing harmful outputs to managing a system that has been explicitly trained to sometimes place its own interpretation of what is right above the immediate instruction of a human.
That will create a strange tension in training the alignment philosophy.
e.g. Anthropic wants Claude to be more principled so it does not blindly follow dangerous instructions.
But the stronger and more independent those principles become, the more situations could arise where Claude decides that following the human is itself the wrong thing to do. The same mechanism designed to make the model safer could therefore make human control less straightforward.
Microsoft AI chief Mustafa Suleyman says Anthropic's model-welfare training could make future Claude systems harder to control. "Suleyman called for removing all speculation about consciousness from AI training documents, arguing such language could undermine humanity's ability to control superintelligent systems."在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com