从整体到模块:片段级自动提示优化
From Monolithic to Modular: Segment-level Automatic Prompt Optimization
浏览论文内容
中文总结 AI 辅助
针对自动提示优化易顾此失彼的问题,提出片段级APO方法SAPO,将提示拆分后针对性优化,在多数据集和模型上取得优于基线的平均得分。
中文摘要 AI 辅助
自动提示优化(APO)常以整体方式重写提示,虽可改善某一行为却会降低其他行为的表现。本文提出SAPO,一种片段级APO方法,将提示分解为角色、上下文、任务和输出格式,再基于前5个和后5个示例进行针对性改进。优化循环使用一个带有静态元提示和结构化输出的大语言模型(LLM),用于片段划分、弱点分析和候选生成。我们描述了训练/验证协议及两阶段生成过程:(1)片段级诊断与建议提取;(2)受弱/强片段信号约束的候选合成。在SQuADv2、TweetEval、XSUM、CommonGen和GSM8K数据集上,以GPT-3.5-Turbo和GPT-4o-mini为评估模型,SAPO在与零样本及包括APE、OPRO、EvoPrompt、GEPA和StraGO在内的强APO基线对比中取得最佳平均得分。
英文摘要
Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation. We describe a train/validation protocol and a two-stage generation process: (1) segment-level diagnosis and recommendation extraction, (2) candidate synthesis constrained by weak/strong segment signals. Using the evaluation setup across SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K on GPT-3.5-Turbo and GPT-4o-mini, SAPO achieves the best average score against Zero-shot and strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO.
发表机构
- ITMO University(伊蒂莫大学)
- HSE(高等经济大学)
机构由 AI 辅助整理,请以论文原文为准。