跳到正文
@OpenAIDevs· @OpenAIDevs · X·· 14 天前AI 评分29
AI 导读

在衡量 AI 模型能否跨应用完成真实业务工作流的 AutomationBench 上,GPT-6 Sol(xhigh)领先 Fable 5.1(max 搭配 Opus 5 fallback),每任务报告成本低 88%。https://t.co/6NrsfCuzBj

正文

On AutomationBench, which measures whether AI models can complete real business workflows across apps, GPT-6 Sol (xhigh) leads Fable 5.1 (max with Opus 5 fallback) with 88% lower reported cost per task. https://t.co/6NrsfCuzBj

来源:@OpenAIDevs · x.com