跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-02精选AI 评分70
AI 导读

据 The Information 报道,OpenAI 的 Astra 模型据称采用了循环深度(looped transformer)架构,同一批 Transformer 层会在生成下一个 token 前反复处理同一信息。

推荐理由

借 The Information 的报道说清 looped transformer 如何在不增参数的前提下提高单 token 算力,以及它对推理可监控性的影响。

正文

The information reports Astra reportedly uses "recurrent depth," or a "looped transformer," which helped its performance while making some internal reasoning less readable.

A looped transformer, or recurrent-depth model, can run the same information through the same transformer layers multiple times before producing the next token, instead of passing through each layer just once in a fixed stack.

That gives the model more computation per token without proportionally increasing its parameter count, potentially letting a smaller model behave more like a larger one while using less memory and bandwidth.

The concern is that more of this reasoning can happen inside internal numerical states rather than readable chain-of-thought text, making human monitoring harder.

引用@rohanpaul_ai@rohanpaul_ai
OpenAI says Astra is its first model to reach the Critical cybersecurity capability threshold. Under its Preparedness Framework, that means Astra can, with the right tools and access, find unknown flaws and develop exploits across hardened systems without step-by-step human guidance. Hence, OpenAI now says Astra will launch with additional chain-of-thought monitoring, while classifiers can automatically stop potentially unauthorized actions. Astra reached roughly 39% exploit success at ~75K output tokens, while GPT-5.6 Sol is only around 1% there and needs nearly 140K tokens to reach ~12% i.e. Astra is dramatically more capable and token-efficient at exploit development on this internal benchmark. Expert assessments went further: Astra escaped a browser sandbox, executed host commands, and escalated an unprivileged operating-system user to root.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com