arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEVE:通过错误驱动精炼与正则化验证稳定文本梯度提示优化

STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification

Yifan Xu, Yixuan Li, Xinzhuo Li, Yixin Gu, Yifan Shen, Lijun Yu, Haohan Wang

arXiv 2609.23716首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Google DeepMind(伊利诺伊大学厄巴纳-香槟分校; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

STEVE通过错误驱动精炼和正则化验证两种机制,稳定文本梯度提示优化,减少性能退化,在多个基准上生成更稳健的提示。

AI 中文摘要

文本梯度方法通过自然语言反馈自动化提示优化,但其迭代更新可能不稳定。我们识别出这种不稳定性的两个来源:从已正确处理的示例中产生的噪声梯度,以及对困难案例的过度专门化,这会降低在更简单输入上的性能。我们引入STEVE,一个具有两个耦合机制的稳定化框架。错误驱动精炼仅从错误处理的示例中生成梯度,将更新集中在信息丰富的失败案例上。正则化验证将每次更新视为临时的,仅当在困难案例上的改进不会导致在保留集上出现不可接受的性能回退时,才接受该更新。在十个推理基准、三个评估器/优化器模型以及已建立的提示优化基线上,STEVE减少了性能退化并生成了更稳健的提示。使用gpt-5.4-mini/gpt-5.4在符号推理、GSM8K-Platinum和DS-1000上的额外评估表明,这些收益在更新的模型和更大的测试集上依然存在。因此,STEVE提供了一种实用的方式来提高文本梯度提示优化的稳定性和有效性。

英文摘要

Textual-gradient methods automate prompt optimization through natural-language feedback, but their iterative updates can be unstable. We identify two sources of this instability: noisy gradients produced from already-correct examples and over-specialization to hard cases that degrades performance on simpler inputs. We introduce STEVE, a stabilization framework with two coupled mechanisms. Error-Driven Refinement generates gradients only from incorrectly handled examples, concentrating updates on informative failures. Regularized Verification treats every update as provisional and accepts it only when improvement on hard cases does not cause unacceptable regression on a preservation set. Across ten reasoning benchmarks, three evaluator/optimizer models, and established prompt-optimization baselines, STEVE reduces degradation and produces more robust prompts. Additional evaluations with gpt-5.4-mini/gpt-5.4 on symbolic reasoning, GSM8K-Platinum, and DS-1000 show that these gains persist with newer models and larger test sets. STEVE therefore provides a practical way to improve the stability and effectiveness of textual-gradient prompt optimization.

CommentsAccepted to Findings of the Association for Computational Linguistics: AACL-IJCNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑