跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-18AI 评分36
AI 导读

Meta FAIR 提出 Research Preference Model(RPM),在消耗 GPU 前预测哪个候选实验值得执行。AIRA-dojo 每步生成 15 个候选改动,RPM 结合代码与历史结果排序,只放 1 个进入完整运行;推理版 RPM 用冻结 LLM,智能体版 RPM 可先花 5 分钟做试点实验。

正文

AI research agents can generate experiment ideas faster than they can afford to run them.

This Meta FAIR paper attacks that bottleneck: before spending GPU time, a Research Preference Model (RPM) predicts which candidate is worth executing.

AIRA-dojo generates 15 candidate changes at each step.

The RPM sees their code plus prior results, ranks them, and sends only 1 into the expensive full run.

The inference-only RPM uses a frozen LLM.

The agentic RPM can first spend a 5-minute budget on pilot experiments for extra evidence.

Across 20 AIRS-Bench tasks, average normalized score rose from 0.684 with random selection to 0.711 with inference-only RPM and 0.729 with the agentic RPM.

Both RPM variants matched the baseline's 24-hour score in roughly 15 hours.

– arxiv. org/abs/2608.13940

Title: "AI Research Preference Models"

来源:@rohanpaul_ai · x.com