论文测试的 NPO 只保留 1 条提示词谱系,让教师模型根据近期 rollout 轨迹和奖励进行改写,在 IFBench 和 HotpotQA 上以略少的 rollout 次数(3,500 vs 3,593、6,800 vs 6,871)达到与维护多候选并做 Pareto 选择的 GEPA 相当或更好的结果。
Prompt optimization may not need a search tree at all.
It may depend more on the quality of the teacher and its feedback than on how elaborate the search algorithm is.
This paper tests NPO, which keeps 1 prompt lineage and asks a teacher model to revise it from recent rollout traces and rewards.
Against GEPA, which maintains multiple prompt candidates with Pareto-based selection, NPO reached comparable or better results on IFBench and HotpotQA with slightly fewer rollouts: 3,500 vs 3,593 and 6,800 vs 6,871.
The gap widened as the teacher got stronger.
The gap widened as the teacher got stronger.
With DeepSeek-V4-Flash and especially GPT-5.5, NPO improved more consistently than GEPA, suggesting that stronger teacher reasoning plus rich feedback can replace some optimizer-side search.
Those optimized prompts also transferred to other student models, especially within the same model family.
来源:@rohanpaul_ai · x.com