Microsoft 在一篇论文中提出 ThinkingBox,一个让智能体对接真实工具、模拟客户和实时后端,并在事后检查数据库的沙盒与基准。推文指出一次性成功不等于可靠,最好的智能体至少成功过一次的业务任务占 91%,但每次都能成功的只有 25%。5 次失败运行中有 4 次礼貌收尾并调用了写数据库的工具,只从回复上看不出任务并未完成,因此评测改为检查数据库状态。
Very relevant Microsoft paper on agent reliability.
Succeeding once and being reliable are not the same thing.
The best agent solved 91% of business tasks at least once but only 25% every time, so measure repeats.
The failures are also hard to spot from the outside.
4 out of 5 failed runs ended politely and called a tool that writes to the database.
The agent said the job was done, but the records said otherwise.
So Microsoft built ThinkingBox, a sandbox that runs an agent against real tools, a simulated customer, and a live backend, then checks the database afterward instead of reading the reply.
---
– arxiv. org/abs/2608.19741
Title: "One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows"
来源:@rohanpaul_ai · x.com