发表机构
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen); Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK(SZ)(香港中文大学(深圳)理工学院; 深圳市未来智联网络研究院; 广东省未来智联网络重点实验室(香港中文大学(深圳)))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对紧张预算下的自动提示优化,提出单阶段优化器BudgetAPO,采用噪声自适应评估切片、固定配对比较和反思性重写,在七个基准和五个模型上全面超越基线,显著降低调用成本。
AI 中文摘要
自动提示优化(APO)已被广泛用于在不更新权重的情况下适配大型语言模型,并取得了令人鼓舞的结果。然而,现有的方法如GEPA和OPRO假设有数百到数千次主体模型调用,这远远超出了付费、限速API背后的实际可行性。在紧张的预算下,它们以两种方式失败:多阶段流水线可能耗尽预算并返回未改变的种子提示,而单阶段方法则在固定大小的小批量上比较候选方案,而不考虑每个任务的噪声。作为补救措施,我们引入了BudgetAPO,一种适用于紧张预算机制的单阶段优化器。BudgetAPO包含:(1)一种噪声自适应规则,根据通过简短探针测量的每个任务的噪声来调整评估切片的大小;(2)一个固定切片,将每个接受/拒绝决策转化为配对比较;(3)一个反思性操作符,联合重写推理策略和输出格式。在七个基准和五个主体模型上的广泛结果表明,BudgetAPO在每个主体上均排名第一,并在Holm校正的配对检验下击败所有基线,同时在250次调用中返回种子的比例为13%,而GEPA为86%。在GPT-OSS-20B上,GEPA需要4.5倍的调用次数才能达到BudgetAPO在100次调用时的分数。
英文摘要
Automatic prompt optimization (APO) has been widely employed to adapt large language models without updating their weights, yielding promising results. However, existing methods such as GEPA and OPRO assume hundreds to thousands of subject-model calls, far more than is practical behind paid, rate-limited APIs. Under tight budgets they fail in two ways: multi-stage pipelines can exhaust the budget and return the seed prompt unchanged, while single-stage methods compare candidates on fixed-size minibatches, regardless of each task's noise. As a remedy, we introduce BudgetAPO, a single-stage optimizer for the tight-budget regime. BudgetAPO incorporates (1) a noise-adaptive rule that sizes the evaluation slice to each task's noise, measured by a short probe; (2) a fixed slice that turns every accept/reject decision into a paired comparison; and (3) a reflective operator that rewrites reasoning strategy and output format jointly. Extensive results across seven benchmarks and five subject models demonstrate that BudgetAPO ranks first on every subject and beats every baseline under Holm-corrected paired tests, while returning the seed in 13% of runs at 250 calls against 86% for GEPA. On GPT-OSS-20B, GEPA needs 5 times as many calls to match \method's 100-call score.