跳到正文
@AIatMeta· @AIatMeta · X·· 2026-08-21AI 评分31
AI 导读

Meta 预览内部评测框架 WildArtifactBench,用于评估智能体在多样交付格式下的复杂真实任务,采用人类与智能体偏好评判的胜率和 Elo 分数而非严格标准答案评分,以扩展多模态实际工作流的任务覆盖。Meta 已从该基准中开放 10 项任务,用于衡量多模态智能体的实际效用。

正文

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats.

By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows.

We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents: https://t.co/UIpdcw8JF2

来源:@AIatMeta · x.com