LoopArena 将 Qwen3.7-Plus 固定为编码 Worker,只替换决定其下一步任务的 Controller,以隔离智能体的管理能力问题。在全部 27 项任务上,最佳 Controller GPT-5.5 的严格成功率仅 24.69%,而每轮重复原始目标得 18.52%,与 Worker 无控制运行完全相同。
A strong coding model is not enough if the model managing its work does not know when to redirect, verify, or stop.
LoopArena isolates that management problem by fixing Qwen3.7-Plus as the coding Worker and changing only the Controller that decides the Worker’s next assignment.
On full 27-task runs, even the best Controller, GPT-5.5, reached just 24.69% Strict Success Rate; simply restating the original goal every round scored 18.52%, exactly the same as letting the Worker run without control.
useful control has to react to the evolving evidence, shifting the Worker between implementation, verification, recovery, and stopping rather than repeatedly saying “keep going.”
So when evaluating agent systems, benchmark the model that manages the loop separately from the model that writes the code.
– arxiv. org/abs/2608.28281
Title: "LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering"
来源:@rohanpaul_ai · x.com