Hugging Face 联合 Baseten 旗下 Baselabs 与 Goodfire AI,共同推进开源模型的安全与可解释性研究。Thomas Wolf 认为开源不等于无人监管,开源模型反而为更广泛的安全研究提供了更多工具与可见性,Linux 就是开源与安全部署相互促进的例证。更多细节将很快公布。
Happy to start collaborating with @baselabs from @baseten and @GoodfireAI to push safety and interpretability for open models.
As open-source models catch up in performance to frontier models, the community at large have a great opportunity to establish a common practice for effective safety, security and interpretability research.
Some past discussions confused “open” with “unmonitored”. On the contrary, an open model provides much more tooling and visibility for safety research from a broader audience, which has led to a lot of the safety and security techniques we use today.
Linux is a great example of this: open-source and secure deployment are not only compatible but heavily intertwined, as a properly secure system needs an extensive feedback loop of finding and fixing vulnerabilities.
We're excited to share more soon.
来源:@Thom_Wolf · x.com