跳到正文
Arena.ai· @arena · X·· 4 天前精选AI 评分72
AI 导读

Arena 宣布 GPT-6.1 Sol (Max) 以 +11.23% 进入 Agent Arena 第5名,重塑 Pareto 前沿。中位任务成本 $0.56,比 GPT-6 Sol 低 39% 且得分高 1.52 分,比 GPT-6 Astra 低 81% 且差距仅 1.04 分;相比排名第一的 Claude Fable 5.1 (Max) 成本低 88%,差距 3.08 分。

推荐理由

榜单方给出 GPT-6.1 Sol 在 Agent Arena 的排名与成本对比数据,读者可据此评估其性价比位置。

正文

Exciting news: GPT-6.1 Sol (Max) by @OpenAi just landed in the Agent Arena at #5 (+11.23%) and reshaped the Pareto frontier!

At a $0.56 median cost per task, it delivers performance within 2 percentage points of GPT-6 Sol and GPT-6 Astra for substantially less cost:
- 39% lower cost than GPT-6 Sol, while scoring +1.52 pts higher
- 81% lower cost than GPT-6 Astra, while landing within 1.04 pts

GPT-6.1 Sol also delivers top-five performance at substantially lower cost compared to:
- 88% lower cost than Claude Fable 5.1 (Max), while landing within 3.08 pts (ranked #1)
- 65% lower cost than Claude Opus 5.5 (High), while landing within 2.59 pts (ranked #2)
- 80% lower cost than Claude Sonnet 5.5 (Max), while landing within 1.29 pts (ranked #3)

Congrats to the team @OpenAI on this release!

引用Arena.ai@arena
Exciting news: GPT-6.1 Sol (Max) by @OpenAI just landed the Code Arena: WebDev at #3 with 1759 pts, and at a blended $8/MToken it reshapes the Pareto frontier! GPT-6.1 Sol (Max) marks a clear improvement in cost efficiency: it gained 70 points over GPT-6 Sol (Max) for the same price. It landed within 30 points of GPT-6 Astra (Max) at 80% lower blended token cost, and 59 points from Claude Opus 5.5 (Max) at 50% of the price. See position on the Pareto frontier for the Code Arena: WebDev in the post below. Overall, GPT-6.1 Sol improved from GPT-6 Sol by 4 rankings! It also improved in every category: - Consumer Product: #5 → #1 - Simulations: #6 → #3 - Data & Analytics: #4 → #3 - Content Creation Tools: #4 → #3 - Gaming: #6 → #4 - Reference-Based Design: #6 → #4 - Brand & Marketing: #10 → #6 Congrats to the @OpenAI team on the release!
在 X 查看被引用的帖子

来源:Arena.ai · x.com