arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28067cs.AI

SEPO:基于证据的结构化编辑提示优化

SEPO: Evidence-Grounded Prompt Optimization via Structural Editing

Xiaoyu Ma, Haoyue Liu, Yiwen Li, Jionghao Zhu, Zhichao Wang, Ye Chen, Xiaoying Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SEPO提示优化器,通过结构化编辑与编辑-效果谱系反馈优化提示,在14任务套件中较GEPA提升准确率,同时在帕累托前沿上优化效率更高、提示更短。

中文摘要 AI 辅助

现有仅使用API的提示优化器常被描述为可解释的,但实际上这通常仅指事后可检查性:每次迭代仍将提示重写为一个不透明字符串,留下完整提示差异的痕迹,而非可定位的机器可读编辑。本文介绍SEPO(Structural, Evidence-grounded Prompt Optimization,基于证据的结构化提示优化),这是一种以编辑-效果谱系反馈为核心的多轨迹提示优化器。SEPO未将每次迭代视为孤立的完整提示重写,而是在两层提示模式中局部编辑稳定的类型化单元,将每次编辑的目标与实现的结构操作关联到其新修复或破坏的示例,并将此编辑-效果记录向前传递,以指导同一搜索分支上的后续架构调用。这使得提示优化具有可寻址性、可归因性和可操作性。在包含14个任务的保留套件中,SEPO在Llama-3.1-8B-Instruct上比最强基线GEPA提升3.1个百分点,在Qwen3-8B上提升2.2个百分点,分别达到61.9%和73.3%的宏观准确率。SEPO同时处于优化时间和测试时间的帕累托前沿,优化时花费290万优化令牌,而GEPA为410万,生成的提示长度缩短5倍以上。

英文摘要

Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt as one opaque string, leaving a trace of full-prompt diffs rather than localisable, machine-readable edits. This paper introduces SEPO (Structural, Evidence-grounded Prompt Optimization), a multi-trajectory prompt optimiser centred on edit-effect lineage feedback. Rather than treating each iteration as an isolated whole-prompt rewrite, SEPO locally edits stable, typed units in a two-layer prompt schema, links the target and realised structural operations of each edit to the examples it newly fixes or breaks, and carries this edit-effect record forward to guide later architect calls on the same search branch. This makes prompt optimisation addressable, attributable, and actionable. Across a 14-task held-out suite, SEPO improves over the strongest baseline, GEPA, by 3.1 pp on Llama-3.1-8B-Instruct and 2.2 pp on Qwen3-8B, reaching 61.9% and 73.3% macro accuracy. SEPO also lies on both the optimisation-time and test-time Pareto frontiers, spending 2.9M optimisation tokens versus 4.1M for GEPA and producing prompts over 5x shorter.

发表机构

  • School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)理工学院)
  • Xi’an Jiaotong University(西安交通大学)
  • Shenzhen Future Network of Intelligence Institute (FNiI-Shenzhen)(深圳未来智能网络研究院)
  • Guangdong Provincial Key Laboratory of Future Networks of Intelligence, CUHK(SZ)(香港中文大学(深圳)广东省未来网络智能重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑