跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-09-03AI 评分46
AI 导读

FM-Bench 让 15 个前沿模型在 20 个模拟赛季中管理一家足球俱乐部,经历约 340 至 400 个决策点,涉及转会、合同、现金与对手行动。第 5 年的排名与最终结果相关性仅 0.19,DeepSeek-V4-Pro 在第 5、10 年领先却最终排第 12。

正文

Current agent benchmarks may be ending before the real failures start.

FM-Bench shows that a model can look strong after 5 years and still finish far behind after 20, so long-running agents need long-running evaluations.

FM-Bench has 15 frontier models manage a football club for 20 simulated years, across roughly 340 to 400 decision stops where transfers, contracts, cash, investments, and rival actions keep changing the future.

The rankings barely resemble their final shape early on. On seed 1, the year-5 ranking correlated just 0.19 with the final order, and DeepSeek-V4-Pro led at years 5 and 10 but finished 12th.

Competition changes the picture too. In the shared Arena, 10 different models won the league at least once instead of one early leader simply compounding forever.

What tracked stronger performance was managerial behavior: cutting slow-payoff investments near the end, keeping cash deployed, and renewing contracts earlier. Token use spanned about 7X and still did not order the board.

So for long-running agents, short task success is a weak proxy for sustained decision quality. The caveat is that the solo board uses 3 seeds and the Arena only 1 shared world.

– arxiv. org/abs/2608.18423

Title: "FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents"

来源:@rohanpaul_ai · x.com