朴素提示优化:重新思考复杂提示搜索的必要性
Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search
浏览论文内容
中文总结 AI 辅助
该研究提出轻量级朴素提示优化(NPO),发现其性能优于复杂优化器GEPA,且优化后的提示可迁移至同系列其他模型,表明简单线性提示优化可媲美复杂搜索方法。
中文摘要 AI 辅助
在智能体AI中,跨多样任务高效改进自主智能体是加速递归自我改进(RSI)的核心,提示优化作为一种有前景的方法,可带来与微调模型权重相当的性能提升,同时降低优化和推理阶段的计算成本。然而,近期的进展越来越倾向于使用不必要的复杂提示优化器。我们提出朴素提示优化(Naive Prompt Optimization, NPO),这是一种轻量级的单谱系方法,利用带有rollout反馈的教师模型迭代修正提示。NPO使用更少的rollout即可达到与GEPA相当或更优的性能,且随着教师模型性能的提升,其优势会增大,这表明更强的教师推理能力可部分替代优化器侧的搜索复杂度。在交互式游戏中,NPO仍与GEPA大体相当,而GRPO在一些不太适合提示优化的任务上表现更好。我们还发现,经NPO优化的提示直接应用于其他学生模型时,尤其是同一系列的模型,也能带来相近的性能提升。总体而言,我们的初步结果显示,简单的线性提示优化可与复杂得多的搜索程序相媲美。
英文摘要
Efficiently improving autonomous agents across diverse tasks is central to accelerating recursive self-improvement (RSI) in agentic AI, with prompt optimization emerging as a promising approach capable of delivering performance gains comparable to those achieved by fine-tuning model weights, while reducing computational costs in both optimization and serving. However, recent developments increasingly favor unnecessarily complex prompt optimizers. We introduce Naive Prompt Optimization (NPO), a lightweight single-lineage method that iteratively revises prompts using a teacher model with rollout feedback. NPO achieves comparable or better performance than GEPA with fewer rollouts, and its advantage increases with stronger teacher models, suggesting that stronger teacher reasoning can partially substitute for optimizer-side search complexity. In interactive games, NPO remains broadly competitive with GEPA, while GRPO performs better on some tasks less amenable to prompt optimization. We also show that NPO-optimized prompts elicit similar performance improvements when applied verbatim to other student models, especially across models within the same family. Overall, our preliminary results show that simple, linear prompt optimization can rival substantially more sophisticated and complex search procedures.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。