AI 导读
Arena 宣布 Anthropic 的 Claude Haiku 5.5 已上线 Arena 平台,用户可在 Agent Arena 测试,分数即将公布。Agent Arena 基于全球用户贡献的数百万真实长程智能体任务评测模型,模型可使用网页搜索、文件系统和终端工具完成复杂工作流,排行榜以因果追踪方法衡量模型相对平均模型的结果表现。
正文
Claude Haiku 5.5 by @AnthropicAI is now in the Arena. Head to Agent Arena to test it out, scores coming soon!
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Claude Haiku 5.5 is also available in Code Arena: WebDev, Text, Document and Vision.
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released. On average, it costs around 75% less to run than Claude Haiku 4.5.在 X 查看被引用的帖子
来源:Arena.ai · x.com