跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-06AI 评分56
AI 导读

论文提出 Harness-of-Harness(HoH)框架,通过将代码、QA 证据与更新计划带入下一轮编码运行,让编码智能体在多日软件开发中持续改进。在 Codex + GPT-5.5 的 GameCraft-Bench 对比中,3 轮后 HoH 得分 71.52,单纯继续同一编码智能体为 58.24;HoH 在 3 个软件基准上改进了全部 3 种测试的智能体设置。

正文

Harness-of-Harness beat repeated coding-agent runs by carrying code, QA evidence, and plans forward.

Says long-horizon coding is not just about giving an agent more time; it needs persistent project state, independent testing, and replanning from real failures.

The problem: over a long project, coding agents can forget earlier decisions, repeat work, break working features, and miss unfinished requirements.

Harness-of-Harness, or HoH, fixes this by carrying the current software, test evidence, known problems, and an updated plan into the next coding run.

It improved all 3 tested agent setups across 3 software benchmarks.

The clearest comparison used Codex + GPT-5.5 on GameCraft-Bench: after 3 passes, HoH scored 71.52, while simply continuing the same coding agent scored 58.24.

The paper also ran HoH for 70 loops and produced a playable FPS from high-level requirements.

That longer test is only 1 game project and used extra tools and skills, so broader real-world generalization is still open.

So for long-running coding agents, invest in persistent project state, independent QA, and evidence-driven replanning—not just more calls.

来源:@rohanpaul_ai · x.com