斯坦福等机构的新论文提出 DuMateBench 基准,用 200 个从真实用户会话重建的任务评测自主智能体,任务混合编码、网页研究、文档工作和内容创作,环境还包含缺失依赖、不稳定网络和干扰文件。
New Stanford and other top research lab paper shows that the framework around an LLM can change agent performance dramatically.
A strong LLM does not guarantee a strong agent
Their DuMateBench benchmark uses 200 tasks rebuilt from real user sessions.
They mix things agents actually do together: coding, web research, document work, and content creation. The environment also includes missing dependencies, flaky networks, and distracting files.
The clearest result is how much the same model changes across agent frameworks. With Opus-4.8, the final score ranges from 0.5821 with OpenClaw to 0.8548 with DuMate, a 27.27 percentage-point gap.
– arxiv. org/abs/2608.26546
Title: "DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows"
来源:@rohanpaul_ai · x.com