跳到正文
机器之心 · 微信公众号·· 2026-08-22AI 评分53

微软与南京大学提出 LoopsBench,评测长周期 Coding Agent 执行过程

LoopsBench:当Coding Agent开始长期工作,我们该如何重新评测它?

阅读原文

本站未展示全文,请前往来源网站阅读。

AI 导读

微软、南京大学等机构的研究人员提出 LoopsBench,一个面向 long-horizon software engineering 的 Coding Agent Benchmark,将长周期任务拆分为 Development Unit 并表示为 Dependency DAG,用 Ready Frontier 和 Regression Obligation 观察执行过程。

来源:机器之心 · 微信公众号 · mp.weixin.qq.com