Meta 提出 AI Research Preference Models,用冻结的预训练 LLM 预测哪个候选方案最值得投入 GPU 时间,解决研究智能体"想法多、算力少"的筛选难题。
Super interesting paper from Meta.
Long-horizon research agents are coming.
But one common problem with research agents today is the lack of originality and how to decide what experiments are worth exploring.
An AI research agent can propose far more experiments than it can afford to run, so the problem is not idea generation, it's deciding which candidates get GPU time.
AI Research Preference Models is trained to predict which candidate solution is most promising before any of them execute.
Two variants, both built on frozen pretrained LLMs. An inference-only model reasons over candidate plans, code, and previously executed solutions. An agentic model additionally runs small-scale pilot experiments before committing budget.
Dropped into the AIRA-dojo agent and measured on AIRS-Bench, average normalized score moves from 0.684 to 0.711 and 0.729.
Both variants reach the unguided agent's 24-hour performance in roughly 15 hours, using less than two-thirds of its execution budget, and together set new state of the art on two AIRS-Bench tasks.
Paper: https://t.co/djM6W6qhwU
Chat with Paper: https://t.co/3sKj8bu6l0
来源:@omarsar0 · x.com