AI 导读
Parsewave 发布 AutomationBench Verified,对 Zapier AutomationBench 全部 600 个公开任务的验证器进行独立审计,人工确认 206 个真实 bug 并全部修复。
正文
We need more efforts like this.
Every agent benchmark should audit its verifiers.
Parsewave went through all 600 public tasks in Zapier's AutomationBench.
Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206.
Regarding 1,235 Kimi K3 runs, the fixed verifiers changed 27.9% of the grades.
I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.
AutomationBench Verified is out. We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below.在 X 查看被引用的帖子
来源:elvis · x.com