跳到正文
@thsottiaux· @thsottiaux · X·· 27 天前AI 评分29
AI 导读

我个人很高兴 prompt injection 正在全行业得到解决,而且 Astra 不仅是我们迄今为止能力最强的模型,也是我们迄今为止最对齐的模型。不过还有…… https://t.co/w1gWiDKkOC

正文

I for one am glad prompt injection is getting solved across the industry and that Astra is not only our most capable, but also our most aligned model to date. But also ... https://t.co/w1gWiDKkOC

引用@bcherny@bcherny
I am pleased to see that OpenAI’s new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. Nice work! Evaluating and naming other labs turns out to be a great way to encourage them to train more aligned models. We will continue to do this until other labs pay more attention to safety. This is good for everyone and there is a lot of room left to go! We solved prompt injection in practice for Claude models about two months ago. But prompt injection is a significant security risk no matter what model you use, and it is important that the industry similarly spends more effort to train their models to be resistant to prompt injection, among other elements of model alignment. As models become more capable and central to businesses and economies, the risks only increase. We should be taking them seriously, and doing the right thing for our customers and the world.
在 X 查看被引用的帖子

来源:@thsottiaux · x.com