arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

流线型反思进化:面向任务自适应自我精炼流水线

Streamlined Reflective Evolution for Task-Adaptive Self-Refinement Pipelines

Xiaofan Zhou, Lu Cheng

arXiv 2609.32458首次发表:更新:

AI 中文总结

提出工作流设计智能体(WDA)框架,通过联合进化阶段指令与结构,并采用SPLIT解决冗余,实现任务自适应自我精炼流水线,在五个基准上显著提升性能。

AI 中文摘要

反思性提示优化在不更新模型权重的情况下改进大型语言模型(LLM)系统,但固定架构限制了自我精炼的组织方式。我们引入了工作流设计智能体(WDA),一个用于任务自适应自我精炼流水线的流线型反思进化框架。从最小提示开始,WDA联合进化阶段指令及其顺序结构。在进化过程中,我们发现重复修订会在单个提示中积累冗余指令。在WDA中,我们提出用SPLIT解决此问题,它将冗余指令重新分配到专门的阶段。三示例反思和局部筛选指导选择性搜索,而校准分数指导帕累托准入和回滚无用的尾部更新。生成的流水线是任务自适应的:其指令和深度从任务数据中学习,然后对该任务的所有测试输入固定。我们在五个基准上评估WDA,涵盖知识、数学推理、多跳问答和指令遵循。在Qwen3.5-9B上,WDA平均得分为51.24%,比初始求解器提高8.63个百分点,比无SPLIT变体提高2.60个百分点。在GPT-4.1-mini上,得分为49.00%,相应增益为5.62和3.69个百分点。这些结果支持任务自适应自我精炼作为更广泛的智能体工作流搜索的补充方向。

英文摘要

Reflective prompt optimization improves large language model (LLM) systems without updating model weights, but fixed architectures constrain how self-refinement is organized. We introduce Workflow-Designing Agents (WDA), a framework for streamlined reflective evolution of task-adaptive self-refinement pipelines. Starting from a minimal prompt, WDA jointly evolves stage instructions and their sequential structure. During evolution, we find that repeated revisions can accumulate redundant instructions in a single prompt. In WDA, we propose to address this problem with SPLIT, which redistributes these instructions across specialized stages. Three-example reflection and local screening guide selective search, while calibration scores guide Pareto admission and rollback of unhelpful trailing updates. The resulting pipelines are task-adaptive: their instructions and depth are learned from task data, then fixed for all test inputs within that task. We evaluate WDA on five benchmarks spanning knowledge, mathematical reasoning, multi-hop question answering, and instruction following. On Qwen3.5-9B, WDA achieves an average score of 51.24%, improving over the initial solver by 8.63 percentage points and the variant without SPLIT by 2.60 points. On GPT-4.1-mini, it achieves 49.00%, with corresponding gains of 5.62 and 3.69 points. These results support task-adaptive self-refinement as a complementary direction to broader agentic workflow search.

Comments59 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑