自动演化特定任务优化的提示指南
Automatically Evolving Prompt Guidelines for Task-Specific Optimization
浏览论文内容
中文总结 AI 辅助
研究大语言模型提示指南优化问题,提出AGOPS方法,利用现有任务示例隐含信息,通过优化方案自动演化特定任务指南,提升下游任务性能,弥补提示不明确导致的性能损失。
中文摘要 AI 辅助
为使大语言模型可靠回答用户查询,用户须清晰指定要求、上下文和约束。但实际中用户查询常不明确,现有提示工程指南通用且与任务无关,且人工非系统形成。为此研究提示指南优化问题,即自动生成特定任务指南以助写更好指定的提示。基于现有任务示例常隐含完成不明确查询所需信息,提出AGOPS方法,通过涉及提示LLM编写器、求解器LLM和提示演化的优化方案演化特定任务指南,最大化下游示例有效性。推理时,指南助用户编写明确提示,提升LLM有效性。实验表明,提示不明确会使下游任务性能大幅下降(高达95.3%),现有技术难恢复,遵循AGOPS指南可平均提升性能15.5%至81.7%。
英文摘要
For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints. In practice, however, user queries are often underspecified, forcing models to infer unstated assumptions that may misalign with the actual user intent. Existing prompt engineering guidelines aim to mitigate this issue, they are typically generic and task-agnostic, limiting their practical utility. Additionally, existing guidelines are formed manually and in a non-systematic way. To this end, we study prompt guideline optimization: the problem of automatically generating task-specific guidelines that help write better-specified prompts for a given task and model. Our key observation is that existing (completed) task examples (aka reference answers) often implicitly encode the missing information required to complete underspecified queries, including behavioral constraints, contextual assumptions, and evaluation criteria. We therefore propose AGOPS, an automatic approach that evolves task-specific guidelines via an optimization scheme that involves a prompt LLM writer, a solver LLM and prompt evolution, which maximize downstream effectiveness on a set of examples (user queries with reference answers). At inference time, our guidelines help users write well-specified prompts, boosting the effectiveness of LLMs. We show across mathematical reasoning, medical question answering, and coding tasks, that prompt underspecification leads to major drops (up to 95.3%) in downstream task performance (compared to well-specified prompts) and, perhaps more importantly, that this drop can hardly be recovered by existing prompt optimization techniques. Users following AGOPS guidelines can regain this loss (increasing performance between 15.5 to 81.7% on average) consistently across all benchmarks.