跳到正文
@OpenAIDevs· @OpenAIDevs · X·· 14 天前AI 评分28
AI 导读

在 DeepSWE v1.1 上(该基准测试编码智能体在长周期工程任务上的表现),GPT-6 Sol (max) 几乎追平 Claude Fable 5 (xhigh),而每任务成本低约 80%。 GPT-6 Luna (max) 与 Fable 5 (medium) 相当,每任务成本低 96%。https://t.co/RN7DmjpUfJ

正文

On DeepSWE v1.1, which tests coding agents on long-horizon engineering tasks, GPT-6 Sol (max) nearly matches Claude Fable 5 (xhigh) at ~80% lower cost per task.

GPT-6 Luna (max) is comparable to Fable 5 (medium) at 96% lower cost per task. https://t.co/RN7DmjpUfJ

来源:@OpenAIDevs · x.com