AI 导读
Accio 已开源 CommerceAgentBench,这是一个包含 107 项任务的基准,涵盖采购、商品上架、运营、订单履约和售后。 该基准根据智能体在模拟电商技术栈中留下的变更来评分,例如应用的标签、保存的草稿、发布的商品列表以及返回的物流单号。 该公司在三种 harness 下测试了十三个不同的模型家族。 上限停留在 61.7% 👀
正文
Accio has open-sourced CommerceAgentBench, a 107-task benchmark spanning procurement, product listing, operations, order fulfillment, and after-sales.
The benchmark ranks scores based on the changes an agent leaves behind in a mock commerce stack, such as labels applied, drafts saved, listings published, and shipment IDs returned.
The company tested thirteen different model families across three harnesses.
The ceiling sits at 61.7% 👀
来源:@testingcatalog · x.com