Meta FAIR 提出 Research Preference Model(RPM),在消耗 GPU 前预测哪个候选实验值得执行。AIRA-dojo 每步生成 15 个候选改动,RPM 结合代码与历史结果排序,只放 1 个进入完整运行;推理版 RPM 用冻结 LLM,智能体版 RPM 可先花 5 分钟做试点实验。
AI research agents can generate experiment ideas faster than they can afford to run them.
This Meta FAIR paper attacks that bottleneck: before spending GPU time, a Research Preference Model (RPM) predicts which candidate is worth executing.
AIRA-dojo generates 15 candidate changes at each step.
The RPM sees their code plus prior results, ranks them, and sends only 1 into the expensive full run.
The inference-only RPM uses a frozen LLM.
The agentic RPM can first spend a 5-minute budget on pilot experiments for extra evidence.
Across 20 AIRS-Bench tasks, average normalized score rose from 0.684 with random selection to 0.711 with inference-only RPM and 0.729 with the agentic RPM.
Both RPM variants matched the baseline's 24-hour score in roughly 15 hours.
– arxiv. org/abs/2608.13940
Title: "AI Research Preference Models"
来源:@rohanpaul_ai · x.com