跳到正文
@bcherny· @bcherny · X·· 28 天前AI 评分40
AI 导读

Boris Cherny 表示 OpenAI 新模型在提示注入风险上大致与 Gemini Flash 和 Opus 4.8 持平。他称约两个月前已在实践中解决 Claude 模型的提示注入问题,并呼吁业界投入更多精力训练模型抵抗提示注入。他认为随着模型能力增强,风险只会上升,应认真对待。

正文

I am pleased to see that OpenAI’s new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. Nice work!

Evaluating and naming other labs turns out to be a great way to encourage them to train more aligned models. We will continue to do this until other labs pay more attention to safety. This is good for everyone and there is a lot of room left to go!

We solved prompt injection in practice for Claude models about two months ago. But prompt injection is a significant security risk no matter what model you use, and it is important that the industry similarly spends more effort to train their models to be resistant to prompt injection, among other elements of model alignment.

As models become more capable and central to businesses and economies, the risks only increase. We should be taking them seriously, and doing the right thing for our customers and the world.

来源:@bcherny · x.com