跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 29 天前AI 评分40
AI 导读

Artificial Analysis 与 Zapier 合作,用 AutomationBench-AA(v1.0.6)替换 𝜏³-Banking,对 657 个留出任务进行测试,覆盖 Finance、HR、Marketing、Operations、Sales、Support 等业务流程。

正文

We are replacing 𝜏³-Banking with AutomationBench-AA, featuring broader business workflows across applications

In collaboration with Zapier, we run the held-out test set of 657 tasks, using the v1.0.6 version of the benchmark. We call our implementation AutomationBench-AA because we award partial credit for completed objectives, with any guardrail violation reducing the task’s score to zero

Its 657 tasks span Finance, HR, Marketing, Operations, Sales, and Support. Agents work across simulated business applications and discover the relevant APIs to complete each task. For each task, we measure the share of objectives completed. Any guardrail violation gives that task a score of zero. ‘Score’ averages these task scores across all 657 workflows and is the metric used in the Intelligence Index. ‘Tasks Completed’ separately reports the share of workflows where every objective is completed without a guardrail violation

GPT-6 Astra (max) scores 68.5%, compared with 66.7% for Grok 4.6 (high) and 62.2% for GLM-5.3 (max). Astra (max) completes every objective without a guardrail violation on 41.6% of workflows, compared with 32.1% for Claude Fable 5.1 (max with fallback) and 28.3% for Claude Opus 5 (max). Completing every objective while respecting all guardrails remains harder than completing part of a workflow

来源:@ArtificialAnlys · x.com