跳到正文
DAIR.AI· @dair_ai · X·· 3 小时前AI 评分48
AI 导读

Meta AI 发布 MIRA,将长周期研究智能体拆成两层:外层 meta-reasoner 读取持久研究记录并生成下一次调查的工作指令,新的 executor 负责执行每条指令,决策只发生在工作指令边界处。

正文

Banger paper from Meta AI on research agents that decide what to investigate next.

(bookmark it)

If you run long-horizon research agents, choosing the next investigation is hard to learn, because those decisions are rare in long traces and their effects show up several steps later.

MIRA splits the agent into two.

An outer meta-reasoner reads a persistent research record and writes a work order for the next investigation.

A fresh executor carries out each work order.

Decisions only happen at work-order boundaries, so the authors train a critic at those points to forecast remaining return, then a single actor-critic (MIRA-AC) that both values partial progress and picks the next investigation.

Even without training, the split improves theorem proving and open-ended architecture research.

Trained on the model's own proxy signals, MIRA-AC improves gold scores in all four autoresearch environments.

Paper: https://arxiv.org/abs/2610.02525

Chat with Paper: https://academy.dair.ai/papers/learning-what-to-investigate-next-meta-reasoning-for-long-horizon-research-agent-2610.02525

来源:DAIR.AI · x.com