清华论文提出 AAArena,基于其年度机器人竞赛的 12 款游戏、1920 个历史人类程序作为对手,测试编码智能体在不改模型权重下读规则、选对手、看回放并重写机器人。在测试的 3 款游戏中,详细回放均优于仅胜负反馈:Pacman 机器人靠回放拿到第 1 名,无回放仅第 11 名。但比赛预算增至 3 倍,4 个卡住的机器人仍无一登顶,复杂规则游戏依旧难突破。
New Tsinghua paper finds that AI agents improving game bots from match replays can top human leaderboards, but mostly stall on games with complex rules.
Getting AI to learn a winning game strategy from a limited number of matches is still hard, especially against changing rivals.
They built AAArena from 12 games in Tsinghua's yearly bot-building contest, with 1,920 archived human programs as rivals. A coding agent, with its model weights unchanged, reads the rules, picks opponents, studies replays, and rewrites its bot within a match budget.
Detailed replays beat win/loss-only feedback in all 3 games tested. With replays, a Pacman bot reached rank 1, versus rank 11 without them.
Tripling the match budget did not push any of 4 stuck bots to rank 1.
– arxiv. org/abs/2610.12341
Title: "Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition"
来源:Rohan Paul · x.com