跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-02精选AI 评分71
AI 导读

OpenAI 表示 Astra 是其首个达到 Preparedness Framework 关键网络安全能力阈值的模型,将带额外的 chain-of-thought 监控发布,分类器可自动阻止潜在的未授权操作。

推荐理由

原文给出了 Astra 与 GPT-5.6 Sol 在内部基准上的漏洞利用成功率与 token 效率对比,读者可据此判断这次网络安全能力跃升的幅度。

正文

OpenAI says Astra is its first model to reach the Critical cybersecurity capability threshold.

Under its Preparedness Framework, that means Astra can, with the right tools and access, find unknown flaws and develop exploits across hardened systems without step-by-step human guidance.

Hence, OpenAI now says Astra will launch with additional chain-of-thought monitoring, while classifiers can automatically stop potentially unauthorized actions.

Astra reached roughly 39% exploit success at ~75K output tokens, while GPT-5.6 Sol is only around 1% there and needs nearly 140K tokens to reach ~12%

i.e. Astra is dramatically more capable and token-efficient at exploit development on this internal benchmark.

Expert assessments went further: Astra escaped a browser sandbox, executed host commands, and escalated an unprivileged operating-system user to root.

引用@OpenAI@OpenAI
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible. Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework. We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve. https://t.co/OrrTgdU90K
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com