跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 9 天前精选AI 评分78
AI 导读

Anthropic 发布 Claude 5.5 家族的第二款模型 Claude Sonnet 5.5,在 Terminal-Bench 4.0 上得分 70.6%,高于 Sonnet 5 的 10.3%,价格保持不变。

推荐理由

原文把重点放在每任务成本而非榜单分数,读者可据此判断 Sonnet 5.5 在日常编码工作里的性价比。

正文

Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices.

Overall, 30% cost reduction per-task due to faster speeds and fewer tool calls.

Keeps Sonnet 5's $2/$10 per million input/output tokens, half Opus 5.5's rates.

Its savings instead come from doing less work per job, since Anthropic says fewer tokens and tool calls cut the total cost of a task by up to 30%.

Output also arrives more than 30% faster than Sonnet 5's, making Sonnet 5.5 Anthropic's quickest Sonnet yet.

Users can dial an effort setting, trading longer reasoning and more self-checking for a higher cost per task.

The economics might count for more than the leaderboard. Sonnet 5.5 operating at Low or Medium effort is able to exceed Sonnet 5's top score at about one-tenth the cost per task. On FrontierCode, it says Sonnet 5.5 at High effort scores roughly 10 points above Sonnet 5 at the same setting while costing approximately one-fifteenth as much per task.

Anthropic characterizes Sonnet 5.5 as ideal for comparatively well-defined routine work, such as software debugging, coding, document production, building presentations and spreadsheets, and designing or refining interfaces.

引用@claudeai@claudeai
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work. https://t.co/UvXD8mDTF1
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com