跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-24AI 评分51
AI 导读

MerchantBench 让 LLM 智能体经营一家小型网店并运行一个模拟年,负责选品、定价和现金管理,同时额外记录智能体是否仍在行动。表现最好的智能体经营完一个模拟年后收入约为人类的四分之一,主要原因是中途沉默停摆。材料建议记录智能体在每个时间窗口的行动次数并观察这条曲线,因为不错的最终分数可能掩盖它数月前就已经停止工作。

正文

Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart.

MerchantBench hands an agent a small online store and lets it run for a simulated year, sourcing products, setting prices, managing cash. It also scores something most evals skip: whether the agent is still acting at all.

The best agent finished a simulated year of shopkeeping with about a quarter of what humans earned, mostly by going quiet, so track how often yours still acts.

So log your agent's actions per window and watch that curve. A decent final score can hide an agent that stopped working months ago.

– arxiv. org/abs/2607.28956

Title: "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations"

来源:@rohanpaul_ai · x.com