跳到正文
@kimmonismus· @kimmonismus · X·· 2026-08-26AI 评分40
AI 导读

编码智能体可以通过测试,却仍然没完成任务。 在 SWE Refactor Bench 中,520 次运行里有 340 次完成了迁移审计。88 次通过了所有固定测试。只有 28 次挺过了完整的三阶段评估。 这一差距才是自主全仓库重构的真实现状。https://t.co/xhxbW4jATI

正文

Coding agents can pass tests and still miss the task.

In SWE Refactor Bench, 340 of 520 runs completed the migration audit. 88 passed every fixed test. Only 28 survived the full three-stage evaluation.

That gap is the real state of autonomous, whole-repo refactoring. https://t.co/xhxbW4jATI

来源:@kimmonismus · x.com