跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-04AI 评分63
AI 导读

Epoch AI 的能力指数 ECI 显示,GPT-6 Astra 取得 169 分,比此前最高的 163 分高出 6 分,并在数学、持续学习与游戏解谜基准上各自刷新纪录。Epoch 表示这一跃升仍处于推理时代 ECI 趋势的不确定范围内;该指数把 50 多个基准合成为单一能力尺度,并按模型结果重叠情况估计各基准难度,因此难题上的高分比简单题上的接近满分更具信息量。在长程编码基准 MirrorCode 上,Astra 排名介于 Opus 4.7 与 Fable 5 之间,OpenAI 为其测试提供了预发布访问。

正文

GPT-6 Astra hit a record 169 on Epoch AI's capability index, jumping 6 points past the previous ceiling and setting separate records across math, continual learning, and game puzzles.

This benchmark ECI combines scores from more than 50 benchmarks into one capability scale, reducing dependence on any single saturated test.

Epoch estimates each benchmark's difficulty from overlapping model results, so strong scores on harder tests carry more information than near-perfect scores on easier ones.

引用@EpochAIResearch@EpochAIResearch
GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5. OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com