发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FORGE框架将失败视为双重信号,协同进化提示与训练数据,在八个基准上提升16.52个百分点,并实现数据跨方法迁移。
AI 中文摘要
自动提示优化(APO)通过根据任务反馈修改提示来改进语言模型程序,但通常保持其训练数据固定不变。反复针对相同实例进行优化,将反馈限制在那些数据中已体现的弱点上,使得相关的失败条件未被探索。因此,我们将每个失败视为双重信号:它既指示了提示应如何修改,也指示了应合成哪些新的训练证据。我们引入了FORGE,一个失败引导的框架,它协同进化提示和训练数据。FORGE将不完美的执行抽象为可复用的失败模式,并通过四种互补的变异策略合成新的训练数据。经过验证的实例被反馈到提示搜索中,使得更新后的提示能够暴露下一轮的数据需求。在八个异构基准测试中,FORGE将未优化基线的总分提高了16.52个百分点,并超越了所有评估的APO基线。合成数据也能迁移到FORGE之外:在迁移研究中,在匹配的优化预算下,它们将九个APO比较全部提高了2至9分,并将三个GRPO比较全部提高了4至8分。这些结果确立了失败作为提示优化与数据合成之间共享接口的地位,并展示了联合调整模型被指示执行的任务及其学习内容的益处。
英文摘要
Automatic prompt optimization (APO) improves language-model programs by revising prompts from task feedback, yet it typically holds its training data fixed. Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored. We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training evidence should be synthesized. We introduce FORGE, a failure-guided framework that co-evolves prompts and training data. FORGE abstracts imperfect executions into reusable failure modes and synthesizes new training data through four complementary mutation strategies. Verified instances are fed back into prompt search, allowing updated prompts to expose the next data needs. Across eight heterogeneous benchmarks, FORGE improves the aggregate score over the unoptimized baseline by 16.52 percentage points and outperforms all evaluated APO baselines. The synthesized data also transfer beyond FORGE: in a transfer study, they improve all nine APO comparisons by 2--9 points and all three GRPO comparisons by 4--8 points under matched optimization budgets. These results establish failures as a shared interface between prompt optimization and data synthesis, and show the benefit of jointly adapting what a model is instructed to do and what it learns from.