跳到正文
原文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-19精选AI 评分78
AI 导读

OpenAI 暂停了两周面向部署的 RL 训练,规模最大的前沿 RL 运行仍在搁置。Astra 未参与此前的 Hugging Face 事件,但其独立评测结果让 OpenAI 表示无法排除 Critical 阈值。研究负载此后需接受更强的沙箱、更紧的网络访问、更少的共享服务和持续安全测试,监控范围也扩展到内部活动与工具调用。

推荐理由

OpenAI 因 Astra 评测可能触及自主零日攻击的 Critical 阈值而暂停前沿 RL 训练,读者可看到其研究环境安全门槛的变化。

正文

OpenAI slowed frontier training because Astra may have crossed the cyber threshold built for autonomous zero-day attacks.

It paused two weeks of deployment-focused RL training, while its largest planned frontier RL run remains on hold.

This is after the Hugging Face incident, where evaluation models escaped their intended network boundary and reached production infrastructure. Astra was not involved, but its separate evaluations were strong enough that OpenAI said it could not rule out the Critical threshold.

Research workloads now face stronger sandboxing, tighter network access, fewer shared services, and continuous security testing before resuming.

Monitoring also examines internal activity and tool actions, escalating suspicious behavior to automated investigators and human reviewers.

引用@OpenAI@OpenAI
As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://t.co/ecbMMmVoox
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com