亚马逊新论文提出用 LLM 充当"研究世界模型",在实验跑之前预测训练改动能否带来收益。基于 9 个设置、2653 条真实实验记录测试,同设置历史记录将预测与实际结果的平均排名相关性从 0.506 提升至 0.774(5 个设置);在 OLMo3-100M 上,低推理努力加记录得 0.892,最高努力无记录仅 0.648。
New Amazon paper shows that an LLM can act as a research world model, predicting whether a training change will help before anyone spends GPU time on it.
AI research agents can propose experiments far faster than teams can afford to run them. Choosing what gets GPU time means guessing outcomes in advance.
They used an LLM as a research world model that predicts an experiment's gain before it runs. They tested it on 2,653 real experiment records from 9 setups, from pretraining to inference.
Past records from the same setup raised average ranking correlation with actual results from 0.506 to 0.774 across 5 setups. On the OLMo3-100M setup, adding records at low reasoning effort scored 0.892, while max effort without records reached 0.648.
If you run research agents, log every experiment, including failures, and feed those records to whatever model picks the next run.
来源:Rohan Paul · x.com