跳到正文
@kimmonismus· @kimmonismus · X·· 23 天前AI 评分62
AI 导读

Noam Brown 在 The Information 访谈中表示,递归自我改进是 OpenAI 的首要优先级,且领先幅度不小。他认为预训练与强化学习的效果是乘法而非加法关系,并预计再过一两次模型发布,AI 在选择研究方向和长期工作排序上可能超过自己的判断。

正文

Noam Brown (@polynoamial) gave a very interesting interview on The Information on OpenAI’s priorities and what comes next. Here is the tl;dr:

- Recursive self-improvement is OpenAI’s clear priority: “the number one priority is recursive self-improvement and by a pretty wide margin.” Building models that help develop better models comes first.

- AI could surpass his research intuition within a couple of releases. He wouldn’t be surprised if, “one or two model releases from now,” he concludes: “they’re better than me at that too.” He specifically means choosing research directions and prioritizing long-term work.

- Pretraining and reinforcement learning amplify each other: “the effects of these two are not additive, they’re multiplicative.” Brown expects their combined progress to produce much more powerful models.

- AI-generated math is becoming easier to produce than to verify: “the biggest challenge that we face with our math results is … double-checking with human mathematicians.” Human verification remains a bottleneck.

- OpenAI underestimated agents during the security incident: “we trusted the sandboxes” and “we just underestimated the AIS.” Brown says monitoring was subsequently added to training and evaluation.

- Monitoring reasoning could get harder: “the agents are more effective at controlling their chain of thought.” Brown warns that punishing unwanted thoughts can teach models to hide them.

Development is progressing rapidly, but he is worried about safety.

来源:@kimmonismus · x.com