跳到正文
DAIR.AI· @dair_ai · X·· 3 小时前AI 评分59
AI 导读

论文用同一模型在 SWE-bench Verified 上运行 Claude Code、mini-SWE-agent 和 OpenCode,发现 447 个任务上两者分数相差不到 5 分,45 个任务困难集上换 harness 与重跑一样会翻转 13% 的任务。

正文

Useful paper on what an agent harness changes when the model stays the same.

One takeaway: keep your harness prompt and tools small.

In this study, that can cut costs by up to 3x without lowering accuracy.

The authors ran Claude Code, mini-SWE-agent, and OpenCode with the same model on SWE-bench Verified. Claude Code and mini-SWE-agent scored within 5 points of each other on 447 tasks.

Swapping the harness changed results about as much as rerunning it. On a 45-task hard set, both flipped 13% of tasks.

The cost difference comes from the system prompt and tool schemas, which each harness sends again at every step. The more steps the agent takes, the more times you pay for them.

Paper: https://academy.dair.ai/papers/what-does-a-harness-buy-tokens-mostly-2610.04433

来源:DAIR.AI · x.com