AdaGEPA:用于反思性提示优化的自适应反馈分配
AdaGEPA: Adaptive Feedback Allocation for Reflective Prompt Optimization
浏览论文内容
中文总结 AI 辅助
提出AdaGEPA自适应反馈分配方法,依据提示词性能与任务结构选择示例进行反思性提示优化,在六个基准上以匹配预算取得更高验证分数,并提升预算效率。
中文摘要 AI 辅助
提示优化通过改进提示词来提升语言模型系统在下游任务上的性能。经典方法在任务示例上评估提示词,并利用产生的反馈通过反思来指导提示词的修订。然而,当反馈选择未考虑提示词的弱点时,这些修订可能仅提升所选示例上的性能,而无法带来更广泛的任务改进。为解决此问题,我们提出AdaGEPA,一种自适应反馈分配方法,该方法利用提示词的性能和任务结构来选择用于下一次提示词修订的示例。我们的方法在每次反馈小批量中最多替换一个示例,以针对已识别的弱点,同时保留其余反馈上下文。在六个下游基准的主要实验中,在匹配的展开预算下,AdaGEPA相比非自适应反馈选择取得了更高的平均验证分数。AdaGEPA还在多个任务上更早地找到了高性能的提示词。在初始的Schema-Guided Dialogue (SGD)研究中,其半预算的提示词在联合目标准确率上超过了非自适应基线的全预算提示词,这些提示词应用于搜索期间未见和已见的服务中的新对话。总体而言,我们的研究结果凸显了自适应反馈分配在提升反思性提示优化的有效性和展开预算效率方面的潜力。
英文摘要
Prompt optimization improves the performance of language-model systems on downstream tasks by refining their prompts. Classical methods evaluate prompts on task examples and use the resulting feedback to guide prompt revisions through reflection. However, when feedback selection does not account for the prompt's weaknesses, these revisions may improve performance on selected examples without yielding broader task improvements. To address this issue, we propose AdaGEPA, an adaptive feedback-allocation method that uses the prompt's performance and task structure to select examples for the next prompt revision. Our method replaces at most one example in each feedback minibatch to target an identified weakness while preserving the remaining feedback context. Across our main experiments on six downstream benchmarks, AdaGEPA achieves higher mean validation scores than non-adaptive feedback selection under matched rollout budgets. AdaGEPA also finds high-performing prompts earlier across several tasks. In the initial Schema-Guided Dialogue (SGD) study, its half-budget prompts outperform the non-adaptive baseline's full-budget prompts in joint goal accuracy on new dialogues from services seen and unseen during search. Overall, our findings highlight the potential of adaptive feedback allocation to improve both the effectiveness and rollout-budget efficiency of reflective prompt optimization.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。