Meta 首席 AI 官 Alexandr Wang 发文表示坚信必须投入对齐,并提出四点主张。他认为人们和企业只会使用与自身意图和价值观一致的智能体,各实验室应在训练和部署中建立治理框架、引入外部评估者与独立监督,把大部分算力用于服务人而非递归自我改进的竞赛。他引用的 @finkd 帖子提到,Meta 为安全与安保把 Muse 的发布推迟了数月。
We believe strongly in the necessity to invest into alignment.
1. People and businesses will only use agents that are aligned with their intent and values. If we do not build models aligned with people and businesses, then they will move to more aligned options.
2. Every lab should have a strong governance framework across training and deployment. This should include external evaluators, which are best practice for transparency, and independent oversight on things like safety criteria for model launches.
3. Every lab will need to operate within institutional protections of democratic countries. This means labs face significant liability if their models cause harm. This will push the ecosystem in the right ways.
4. Advances in the field are ultimately downstream of compute and resource allocation. Racing on recursive self-improvement is one of the riskiest pathways for potential loss of control to powerful models. Meta is committing the significant majority of our compute towards serving people rather than racing on RSI, and over labs can choose to do the same.
AI is a very powerful technology, and there is immense responsibility in developing it safely alongside the right checks and balances.
Last month I wrote about how we can build a positive and safe future for everyone: https://t.co/eoLGVY8yad Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.在 X 查看被引用的帖子
来源:@alexandr_wang · x.com