跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 25 天前AI 评分47
AI 导读

Devin Fusion CLI 搭配 Claude Fable 5.1(xhigh)与 SWE-2(medium)在 Artificial Analysis Coding Agent Index v1.5 上得分 61.7。

正文

Devin Fusion performs well for cost efficiency and performance, and currently sits on the Pareto frontier for Coding Agent Index score vs. Cost per Task in both configurations we tested

Devin Fusion CLI with Claude Fable 5.1 (xhigh) + SWE-2 (medium) scores 61.7 on the Artificial Analysis Coding Agent Index v1.5. This is almost tied with Claude Fable 5.1 (max, with fallback) in Claude Code at 62.2 despite the lower effort, and costs 36% less at $7.9 per task for the Fusion configuration vs. $12.4 for Claude Code. Speed is also essentially flat, with time per task of 35.8 vs. 34.8 minutes.

This pattern holds across the underlying evaluations: Fusion scores 63.1 vs. 64.3 on DeepSWE 1.1, 65.9 vs. 64.8 on SWE-Atlas QnA, and 56.1 vs. 57.6 on Terminal-Bench 4.0.

来源:@ArtificialAnlys · x.com