arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32452cs.CLcs.LG

反射器应看到什么?反射式提示优化中证据的实证研究

What Should the Reflector See? An Empirical Study of Evidence in Reflective Prompt Optimization

Xiaofan Zhou, Lu Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过实证比较九种反射策略,发现仅失败样本策略在最终性能上最优,但局部反射成功与最终提示强度无必然关联,且校准差距需谨慎解读。

中文摘要 AI 辅助

反射式提示优化利用模型行为的示例来修订指令,但反射器应接收何种证据仍不清楚。我们在单亲Pareto引导搜索中研究了证据组成、示例可见性、候选选择和领域知识策略。使用Qwen3.5-9B作为任务模型和反射器,我们在五个数据集上评估了九种反射策略。从这些实验中,我们发现了三种不同的模式。对于性能提升,仅失败样本产生了最大的平均测试增益(+8.0个百分点),而平衡混合和无示例+验证集共享最佳平均性能排名。对于反射有效性,无示例在改进采样亲本方面取得了最佳排名,但仅产生1.4个百分点的平均测试增益:局部反射成功并不必然产生更强的最终提示。对于过拟合评估,仅失败样本和平衡混合共享最低的平均校准差距排名,而GPQA和IFBench上较大的差距表明校准增益可能夸大保留集的改进。这一差距是一个描述性指标,而非过拟合的直接度量。综合来看,这些结果说明了为何反射策略应在最终性能、亲本改进和校准到测试的迁移上分别评估。

英文摘要

Reflective prompt optimization revises instructions using examples of a model's behavior, but which evidence the reflector should receive remains unclear. We study evidence composition, visibility of examples, candidate selection and domain-knowledge policy within a single-parent Pareto-guided search. Using Qwen3.5-9B as both task model and reflector, we evaluate nine reflection strategies on five datasets. From these experiments, we find three distinct patterns. For performance improvement, Failures-only produces the largest mean test gain (+8.0 percentage points), while Balanced-mix and No-examples+Val share the best mean performance rank. For reflection effectiveness, No-examples achieves the best rank for improving sampled parents, yet yields only a 1.4-point mean test gain: local reflection success does not necessarily produce a stronger final prompt. For overfitting assessment, Failures-only and Balanced-mix share the lowest mean calibration-gap rank, while larger gaps on GPQA and IFBench show that calibration gains can overstate held-out improvement. This gap is a descriptive indicator, not a direct measure of overfitting. Together, these results show why reflection strategies should be assessed separately on final performance, parent improvement and calibration-to-test transfer.

补充信息

↑