跳到正文
Arena.ai· @arena · X·· 4 天前精选AI 评分65
AI 导读

Arena 发布 Agent Arena 更新,Claude Sonnet 5.5 (Max) 以 +12.5% 净改进分首次位列第 3,比 Claude Sonnet 5 (High) 高 8.1 个百分点,并在 Chat 类目以 +15.6% 排名第 1。其单任务中位成本为 $2.74,比排名第 2 的 Claude Opus 5.5 (High) 的 $1.58 高约 73%,因此略偏离 Agent Arena 的 Pareto 前沿;Anthropic 模型目前包揽 Agent Arena 前三。

推荐理由

榜单方基于自家 Agent Arena 数据,指出 Claude Sonnet 5.5 相对 Opus 5.5 在成本与得分上的取舍,读者可据此比较两者性价比。

正文

Claude Sonnet 5.5 by @AnthropicAI just landed at #3 in the Agent Arena. This model has a median cost per task of $2.74, and a +12.5% net improvement score.

Claude Sonnet 5.5 delivers top-tier performance, but at a cost premium: #2 Claude Opus 5.5 costs $1.58 per task while achieving a higher score. That tradeoff keeps Sonnet 5.5 just off the Agent Arena Pareto frontier.

引用Arena.ai@arena
Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement! This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%). This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58. @AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team!
在 X 查看被引用的帖子

来源:Arena.ai · x.com