亚马逊与微软新论文提出 SPACE,让长程智能体不必每个微小动作都调用 LLM 决策,而是从成功轨迹中学习哪些动作可安全合并执行。该方法将轨迹转为程序化技能,以子技能边界标注有意义的动作块,再蒸馏为可变长度动作策略,测试时无需技能库。在 ScienceWorld 上成功率从 35.9% 升至 67.2%,平均 LLM 轮数从 10.2 降至 5.2。
New Amazon Microsoft paper shows long-horizon agents should not need an LLM decision after every tiny action; the hard part is knowing which actions can safely run together.
SPACE learns those boundaries from successful trajectories.
It converts trajectories into programmatic skills, treats subskill boundaries as labels for meaningful chunks, then distills them into a policy that emits variable-length primitive actions with no skill library at test time.
On ScienceWorld, success rose from 35.9% to 67.2%, while average LLM rounds fell from 10.2 to 5.2.
do not make agents reconsider every tiny step, and do not blindly batch actions either. Train them to learn when to keep acting and when to look again.
来源:@rohanpaul_ai · x.com