AI 导读
在 DeepSWE v1.1 上(该基准测试编码智能体在长周期工程任务上的表现),GPT-6 Sol (max) 几乎追平 Claude Fable 5 (xhigh),而每任务成本低约 80%。 GPT-6 Luna (max) 与 Fable 5 (medium) 相当,每任务成本低 96%。https://t.co/RN7DmjpUfJ
正文
On DeepSWE v1.1, which tests coding agents on long-horizon engineering tasks, GPT-6 Sol (max) nearly matches Claude Fable 5 (xhigh) at ~80% lower cost per task.
GPT-6 Luna (max) is comparable to Fable 5 (medium) at 96% lower cost per task. https://t.co/RN7DmjpUfJ
来源:@OpenAIDevs · x.com