AI 导读
Meta 预览内部评测框架 WildArtifactBench,用于评估智能体在多样交付格式下的复杂真实任务,采用人类与智能体偏好评判的胜率和 Elo 分数而非严格标准答案评分,以扩展多模态实际工作流的任务覆盖。Meta 已从该基准中开放 10 项任务,用于衡量多模态智能体的实际效用。
正文
Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats.
By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows.
We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents: https://t.co/UIpdcw8JF2
来源:@AIatMeta · x.com