跳到正文
原文
Rohan Paul· @rohanpaul_ai · X·· 2 小时前AI 评分54
AI 导读

论文提出 AgentBug-Smith,把真实 GitHub bug 报告自动转化为可运行测试,构建含 200 个可复现 bug 且持续增长的 Live-Harness-Bench。最好的编码智能体只修复 9% 的此类 bug,远低于常规软件 bug 约 40% 的水平;用过往修复经验提炼的简短指南后,一个智能体在 79 个未见 bug 上的正确修复从 1 个升到 6 个。

正文

Self-improving AI agents will need to fix their own code.

And this paper from top US+China labs, shows coding agents miss most such bugs but improve with lessons from past fixes.

that real bugs in agent harnesses, can be automatically turned into a growing set of runnable tests.

An agent's own code is everything around the model: tool calls, memory, and prompts. Its bugs depend on live model calls, which makes them hard to recreate and test.

So the researchers built AgentBug-Smith, which turns real GitHub bug reports into runnable tests. The result is a 200-bug benchmark that keeps growing.

The best of 3 coding agents fixed just 9% of those bugs, versus about 40% reported on regular software bugs. A short guide of lessons from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs.

Before trusting a coding agent with your agent's code, try it on bugs you've already fixed.

来源:Rohan Paul · x.com