MIT 等机构论文研究了拍卖与匹配环境中交互格式对 LLM 智能体决策的影响,发现简化呈现和陈述规则事实比要求更多推理更有效。GPT-4o、Claude、Gemini 和 Gemma 仍倾向压低出价以保留利润;升价加留或走的简单选择使 Gemma 的出价差距从低于价值 $5.30 缩到 $0.30,一行说明拒绝只会重新分配的注释将匹配错误从 4.2% 降到 0.2%,而书面计划未体现这些改进。
New MIT Paper: LLM agents decide better when the choice is shown in simple steps or the rule's safe move is stated plainly, and worse when told to reason about opponents.
Market-design rules of thumb built for human bidders carry over to LLM agents, so we can borrow them instead of inventing new prompt tricks.
Honest bids and rankings are always the best move in these auctions and matching games. GPT-4o, Claude, Gemini and Gemma still underbid, often to keep a profit margin.
A rising price with a simple stay-or-exit choice moved Gemma from $5.30 below its value to $0.30 below. A 1-line note that rejections only redirect cut matching errors from 4.2% to 0.2%.
Fix the format and state the key fact before adding reasoning prompts, and judge agents by their choices, since their written plans missed these gains.
Agent prompts should state facts about the rules rather than request more thinking, because facts improved choices while thinking prompts often added errors.
来源:Rohan Paul · x.com