AgentMercury 提出用与评测集无关的合成商业环境训练智能体,仍能提升目标 benchmark 表现。它从纯商业描述生成 4,783 家模拟公司,每家拥有独立服务、工具和数据库,再从中抽取训练任务。环境构建本身也可学习:某模型仅能为 3.3% 的 30 个未见简报生成有效公司,用构建轨迹微调后升至 83.3%,追平 Claude Opus 4.8。
Should your agent's training environment look like your eval set?
AgentMercury says no, and shows that worlds built from business scenarios transfer further.
Train an agent inside a fake company that has nothing to do with your benchmark, and it still gets better at your benchmark.
AgentMercury generated 4,783 simulated companies from plain business descriptions, each with its own services, tools and database, then pulled training tasks out of them afterward.
Building the worlds turned out to be learnable too. One model authored a valid company for only 3.3% of 30 unseen briefs, and 83.3% after fine-tuning on the construction traces, matching Claude Opus 4.8.
– arxiv. org/abs/2608.20634
Title: "AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale"
来源:@rohanpaul_ai · x.com