ASI-Bench 基准给出 11 个科学领域的 60 个真实研究项目,用于测试不同指令粒度对研究型智能体的影响。在 18 组智能体与模型组合中,给出完整步骤时平均分为 50.91,只给方法名时降至 29.10,连方法都不给只再低 2.5 分,说明分数损失主要来自缺少步骤。只给方法名的提示词也最贵,比完整指令多消耗 59% 的 token。
Telling a research AI agent which method to use is close to telling it nothing at all.
What actually carries performance is the procedure, so write the steps rather than the method name.
ASI-Bench gives agents 60 real research projects across 11 scientific fields. Same goal, same data, same scoring every time. Only the instructions change: full procedure, method name only, or nothing but the objective and the data.
Across 18 agent and model combinations, the average score fell from 50.91 with the full procedure to 29.10 with only the method named.
Dropping the method as well cost another 2.5 points. Nearly all the damage comes from losing the steps.
Method-only prompts were also the most expensive to run, burning 59% more tokens than complete instructions and more than prompts that named no method at all. Naming an approach pins the agent to a direction while still leaving it to rebuild every implementation detail.
来源:@rohanpaul_ai · x.com