跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-23AI 评分56
AI 导读

Microsoft 在一篇论文中提出 ThinkingBox,一个让智能体对接真实工具、模拟客户和实时后端,并在事后检查数据库的沙盒与基准。推文指出一次性成功不等于可靠,最好的智能体至少成功过一次的业务任务占 91%,但每次都能成功的只有 25%。5 次失败运行中有 4 次礼貌收尾并调用了写数据库的工具,只从回复上看不出任务并未完成,因此评测改为检查数据库状态。

正文

Very relevant Microsoft paper on agent reliability.

Succeeding once and being reliable are not the same thing.

The best agent solved 91% of business tasks at least once but only 25% every time, so measure repeats.

The failures are also hard to spot from the outside.

4 out of 5 failed runs ended politely and called a tool that writes to the database.

The agent said the job was done, but the records said otherwise.

So Microsoft built ThinkingBox, a sandbox that runs an agent against real tools, a simulated customer, and a live backend, then checks the database afterward instead of reading the reply.

---

– arxiv. org/abs/2608.19741

Title: "One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows"

来源:@rohanpaul_ai · x.com