arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAGE:理解多组件提示优化中的稳定性-性能权衡

MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization

Prateek Singh

arXiv 2607.11944首次发表:更新:

AI 中文总结

研究迭代提示优化组件交互,借助MAGE框架发现提示优化耦合效应,有基于失败反思重要、MAGE性能优、增加候选多样性揭示效应等发现,还验证其与模型余量有关及低数据下固定提示更优,表明应从性能和稳定性评估提示优化系统。

AI 中文摘要

我们通过MAGE(内存增强目标导向提示进化)来研究迭代提示优化的不同组件如何相互作用以及组合时会发生什么,MAGE是一个用于研究提示优化中组件交互的可控分析框架。实验发现了提示优化耦合效应(POCE),即多个随机优化信号在封闭反射回路中运行时,会同时提高性能并放大方差。还得出三个主要发现:基于失败的反思至关重要;MAGE在GSM8K-Hard上比GEPA表现更好;增加候选多样性能揭示最清晰的POCE信号。进一步验证表明POCE与模型余量有关,在低数据情况下,精心设计的固定提示优于所有反射优化器。结果表明提示优化系统应从性能和稳定性两方面评估。

英文摘要

How do different components of iterative prompt optimization interact, and what happens when they are combined? We investigate this through MAGE (Memory-Augmented Goal-directed Prompt Evolution), a controlled analysis framework for studying component interaction in prompt optimization. MAGE is not proposed as a superior optimizer in absolute terms; it integrates episodic memory, multi-objective Pareto selection, and adaptive evaluation as a platform for controlled ablation. Our experiments uncover a previously unreported phenomenon, the Prompt Optimization Coupling Effect (POCE): when multiple stochastic optimization signals operate within a closed reflective loop, they interact in ways that simultaneously improve performance and amplify variance, behavior that cannot be predicted by analyzing components in isolation. Three main findings emerge. First, failure-grounded reflection is essential: methods relying only on scores (OPRO) or abstract critique (Self-Refine) fail to improve prompts. Second, MAGE achieves 46.4% versus GEPA's 34.0% on GSM8K-Hard (+12.4%, P(MAGE>GEPA)=0.998, 5 seeds on gpt-4o-mini), with comparable variance (7.3% vs. 7.0%). Third, increasing candidate diversity reveals the clearest POCE signal: expanding the candidate pool from n=3 to n=5 improves mean accuracy by +21.6% while increasing variance by 3.7x. We further validate on Llama 3.1 8B and show POCE is headroom-dependent: when the base model already achieves high accuracy, variance amplification disappears. Finally, in low-data regimes (Ntrain=30), well-designed fixed prompts outperform all reflective optimizers, indicating that scaffold choice dominates optimizer choice. Our results suggest prompt optimization systems behave as coupled stochastic processes and should be evaluated in terms of both performance and stability, not just peak accuracy.

Comments10 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑