AI 导读
在衡量 AI 模型能否跨应用完成真实业务工作流的 AutomationBench 上,GPT-6 Sol(xhigh)领先 Fable 5.1(max 搭配 Opus 5 fallback),每任务报告成本低 88%。https://t.co/6NrsfCuzBj
正文
On AutomationBench, which measures whether AI models can complete real business workflows across apps, GPT-6 Sol (xhigh) leads Fable 5.1 (max with Opus 5 fallback) with 88% lower reported cost per task. https://t.co/6NrsfCuzBj
来源:@OpenAIDevs · x.com