MerchantBench 让 LLM 智能体经营一家小型网店并运行一个模拟年,负责选品、定价和现金管理,同时额外记录智能体是否仍在行动。表现最好的智能体经营完一个模拟年后收入约为人类的四分之一,主要原因是中途沉默停摆。材料建议记录智能体在每个时间窗口的行动次数并观察这条曲线,因为不错的最终分数可能掩盖它数月前就已经停止工作。
Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart.
MerchantBench hands an agent a small online store and lets it run for a simulated year, sourcing products, setting prices, managing cash. It also scores something most evals skip: whether the agent is still acting at all.
The best agent finished a simulated year of shopkeeping with about a quarter of what humans earned, mostly by going quiet, so track how often yours still acts.
So log your agent's actions per window and watch that curve. A decent final score can hide an agent that stopped working months ago.
– arxiv. org/abs/2607.28956
Title: "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations"
来源:@rohanpaul_ai · x.com