跳到正文
原文
Artificial Analysis· @ArtificialAnlys · X·· 10 小时前精选AI 评分70
AI 导读

Artificial Analysis Coding Agent Index 本周纳入 Claude Sonnet 5.5、GPT-6.1 Sol 和 Gemini 4 Argon,三者均接近榜首但成本差异大。

推荐理由

原文给出三大新模型组合在 Coding Agent Index 的得分与单任务成本对比,读者可以据此在性能和价格之间做选型权衡。

正文

This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index leaderboard, but each has a different balance of performance and cost

The Artificial Analysis Coding Agent Index measures agents (a combination of model and harness) across three agentic coding evaluations.

➤ Claude Sonnet 5.5 (max) in Claude Code takes the top spot at 68, but also has the highest measured cost per task: $14.19

➤ Gemini 4 Argon (high) in Antigravity CLI scores 64 at $5.84 per task - less than half of Sonnet 5.5’s cost. Note, this uses Google’s promotional pricing, and Argon is not yet publicly available

➤ GPT-6.1 Sol (xhigh) in Codex scores 63 at $1.04, roughly one sixth of Argon’s cost

来源:Artificial Analysis · x.com