一篇新论文提出 NPO 提示词优化器,仅保留单一谱系,不做候选池和搜索树,在更小的 rollout 预算下追平 GEPA。它在两个指令遵循基准上分别用 3,500 和 6,800 次 rollout,对比 GEPA 的 3,593 和 6,871 次,并在 22 个 TextArena 游戏中保持大致相当。NPO 的优势随教师模型变强而扩大,暗示优化器侧的搜索复杂度一直在补偿教师推理能力的不足。
Interesting paper on prompt optimization.
They claim that a single-lineage prompt optimizer just matched GEPA on a smaller rollout budget.
Prompt optimization has been drifting toward heavier machinery, with candidate pools, reflection trees, and Pareto-based selection.
NPO keeps one lineage. At each iteration it runs the student on the current prompt, collects rollout traces and rewards, and hands a sliding window of recent iterations to a teacher model that rewrites the prompt.
There is no candidate population and no search tree.
On the two instruction-following benchmarks it spends 3,500 and 6,800 rollouts against GEPA's 3,593 and 6,871, and it stays broadly comparable across 22 TextArena games.
The interaction with teacher strength is what makes this interesting. NPO's advantage grows as the teacher model gets stronger, which suggests optimizer-side search complexity has been compensating for weak teacher reasoning all along.
Paper: https://t.co/rez9FnNZhW
Chat with Paper: https://t.co/sy7HME8Z99
来源:@omarsar0 · x.com