跳到正文
@testingcatalog· @testingcatalog · X·· 2026-09-04AI 评分53
AI 导读

OpenAI 的 GPT-6 Astra 在 ARC-AGI-3 上以标准测试框架取得 63% 分数,动作效率超过人类基线,在 96% 的关卡中使用的动作数少于人类中位数。帖文提到,Astra 的一项关键行为是把陌生环境转化为紧凑的符号化世界模型,用逻辑规则表示游戏机制,并自建领域专用语言记法来跟踪状态、规划动作。

正文

OPENAI 🔥: Astra scored 63% on ARC-AGI-3 with a standard harness.

> GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.

> A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions.

来源:@testingcatalog · x.com