arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04197cs.CLcs.AI

ESPO:基于诊断、多样化与稳定化的错误结构化提示优化

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有进化式提示优化器的提示膨胀问题,ESPO通过诊断、多样化、稳定化三阶段优化,在七个NLP基准及四个学生模型上均提升准确率,缩短提示长度并加快推理。

中文摘要 AI 辅助

GEPA等进化式提示优化器存在提示膨胀问题:每次迭代都会添加规则和注意事项,导致提示长度达到原来的3倍,但准确率并未提升。我们将此问题追溯至三个缺陷:错误观察不完整、搜索多样性有限以及选择不可靠,进而提出ESPO(错误结构化提示优化),将提示优化分解为三个阶段:诊断阶段在一轮内将所有训练错误聚类为结构性模式;提出阶段通过四种具有独立偏差的互补策略生成候选方案;选择阶段采用自助稳定选择。在七个公开NLP基准——Tweet、MMLU、GSM8K、HotpotQA、ScoNe、HoVer和PUPA上,ESPO的平均准确率较当前最优方法提升了3.76个百分点(74.67% vs GEPA的70.91%),在每个数据集上的表现均与GEPA相当或更优,同时生成的提示长度缩短了47%(1004字符 vs 1878字符),推理速度更快。针对另外四个学生模型(Gemma 3 12B、Mistral 14B、Qwen3 32B、Claude Haiku 4.5)的跨模型实验显示,ESPO在所有测试模型上均取得了最佳平均准确率,其中在Qwen3 GSM8K上的差距最大(从15.00%提升至91.40%)。附录中的泛化边界将每个阶段与测试时间差距的对应项关联,消融实验证实了一项关键预测:仅增加多样性而不采用自助选择实际上会损害性能(下降1.20%)。

英文摘要

Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; Propose generates candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection. On seven public NLP benchmarks - Tweet, MMLU, GSM8K, HotpotQA, ScoNe, HoVer, and PUPA - ESPO improves average accuracy by $+$3.76 pp over the state-of-the-art (74.67% vs 70.91% for GEPA), matching or exceeding GEPA on every dataset while producing prompts 47% shorter (1,004 vs 1,878 chars) and faster at inference. Cross-model experiments across four additional student models (Gemma 3 12B, Mistral 14B, Qwen3 32B, Claude Haiku 4.5) show ESPO yields the best average accuracy on every model tested, with the largest gap on Qwen3 GSM8K (15.00% $\to$ 91.40%). A generalization bound (Appendix) grounds each phase in a corresponding term of the test-time gap, and the ablation confirms a key prediction: adding diversity without bootstrap selection actually hurts performance ($-$1.20%).

发表机构

  • AWS Agentic AI(亚马逊云科技智能体人工智能部门)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑