arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11219cs.AIcs.CLcs.LG

从整体到模块:片段级自动提示优化

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

Nikita Kulin, Viktor Zhuravlev, Artur Khairullin, Sergey Muravyov, Ilya Makarov, Daniil Sukhorukov, Ekaterina Averkova

首次发表
浏览论文内容

中文总结 AI 辅助

针对自动提示优化易顾此失彼的问题,提出片段级APO方法SAPO,将提示拆分后针对性优化,在多数据集和模型上取得优于基线的平均得分。

中文摘要 AI 辅助

自动提示优化(APO)常以整体方式重写提示,虽可改善某一行为却会降低其他行为的表现。本文提出SAPO,一种片段级APO方法,将提示分解为角色、上下文、任务和输出格式,再基于前5个和后5个示例进行针对性改进。优化循环使用一个带有静态元提示和结构化输出的大语言模型(LLM),用于片段划分、弱点分析和候选生成。我们描述了训练/验证协议及两阶段生成过程:(1)片段级诊断与建议提取;(2)受弱/强片段信号约束的候选合成。在SQuADv2、TweetEval、XSUM、CommonGen和GSM8K数据集上,以GPT-3.5-Turbo和GPT-4o-mini为评估模型,SAPO在与零样本及包括APE、OPRO、EvoPrompt、GEPA和StraGO在内的强APO基线对比中取得最佳平均得分。

英文摘要

Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation. We describe a train/validation protocol and a two-stage generation process: (1) segment-level diagnosis and recommendation extraction, (2) candidate synthesis constrained by weak/strong segment signals. Using the evaluation setup across SQuADv2, TweetEval, XSUM, CommonGen, and GSM8K on GPT-3.5-Turbo and GPT-4o-mini, SAPO achieves the best average score against Zero-shot and strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO.

发表机构

  • ITMO University(伊蒂莫大学)
  • HSE(高等经济大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑