AI 导读
Astra 在我们的长时程编程基准 MirrorCode 上取得了 46.7% 的原始得分,恰好落在 Opus 4.7 和 Fable 5 之间。 我们进行了一些内部测试,表明 Astra 在更高推理投入下可能在 MirrorCode 上得分更高。我们这里只报告 High-effort 得分,因为这是我们对其他所有模型报告的口径,而且我们还没有在 Max 下运行其他模型。 https://t.co/BVozmfyCDq
正文
Astra achieved a raw score of 46.7% on MirrorCode, our long-horizon coding benchmark landing squarely between Opus 4.7 and Fable 5.
We ran some internal tests indicating that Astra might score higher on MirrorCode with greater reasoning effort. We only report the High-effort score here, as it's what we report for every other model, and we haven't run other models at Max yet.
https://t.co/BVozmfyCDq
来源:@EpochAIResearch · x.com