arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12127cs.CLcs.SE

局部编辑,全局涟漪:基于重放信息的策略适配用于工作流合成

Local Edits, Global Ripples: Replay-Informed Policy Adaptation for Workflow Synthesis

  • Amazon, Inc.(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

Manqing Mao, Hong Wang, Samson Koelle, Jie Yuan, Zhuoer Wang, James Feng, Yanjun Lin, Daniel Edmiston, Nikki Lijing Kuang, Zhecheng Sheng, Wei Niu

AI总结:

针对提示词编辑的局部性与组合敏感性问题,提出RIPPLE方法,分离编辑位置与安全性决策,通过重放评估保留安全编辑,在Flow-HO基准上验证成功率提升达23.1%。

AI中文摘要:

提示词-策略编辑提供了一种实用的方法来改进智能体,使其无需更新底层模型即可合成可执行的工作流。然而,持续的提示词编辑具有两个相互关联的特性。首先,编辑的局部性并不意味着效果的局部性:局限于一个策略片段的编辑可能会波及下游执行,从而改变编辑片段之外的行为。其次,编辑效果对组合敏感:单独有效的编辑在组合后可能相互干扰,导致其中一个或两个编辑失去其益处或变得有害。因此,持续的适配必须支持两个不同的决策:从执行反馈中识别策略应在何处更改,以及确定在组合后所产生的编辑是否仍能安全地持久化。为了应对这些挑战,我们引入了RIPPLE(基于重放信息的持久策略定位与编辑),它将编辑在何处进行与编辑在组合后是否安全这两个问题分离开来。它诊断失败的轨迹,将每个可操作的失败映射到预定义的策略片段,并将修正限制在策略的那一部分。然后,RIPPLE根据相同的迭代起始策略评估候选编辑,以比较其孤立收益,然后在先前接受的更新之后重放有希望的编辑,以揭示下游效应和交互。只有那些在组合下仍然安全的编辑才会被保留。我们在Flow-HO(一个用于可执行工作流合成的合成保留基准)上评估了RIPPLE。RIPPLE将验证成功率提高了最多23.1%,并在两个额外的冻结语言模型骨干上产生了正向收益,同时保持了编辑效率和低执行成本。针对性的交互分析进一步证明了这两个特性:一个片段局部的工具使用编辑改变了下游的资源解析和验证,而一个在孤立情况下有益的编辑在组合后变得有害。

英文摘要:

Prompt-policy editing offers a practical way to improve agents that synthesize executable workflows without updating the underlying model. However, persistent prompt editing has two coupled properties. First, edit locality does not imply effect locality: an edit confined to one policy segment can ripple through downstream execution, altering behavior beyond the edited segment. Second, edit effects are composition-sensitive: edits that work in isolation can interfere after composition, causing one or both to lose their benefit or become harmful. Persistent adaptation must therefore support two distinct decisions: identifying where the policy should change from execution feedback, and determining whether the resulting edit remains safe to persist after composition. To address these challenges, we introduce RIPPLE (Replay-Informed Persistent Policy Localization and Editing), which separates where an edit is made from whether it remains safe after composition. It diagnoses failed trajectories, maps each actionable failure to a predefined policy segment, and restricts the correction to that part of the policy. RIPPLE then evaluates candidates against the same iteration-start policy to compare their isolated gains, before replaying promising edits after previously accepted updates to expose downstream effects and interactions. Only edits that remain safe under composition are retained. We evaluate RIPPLE on Flow-HO, a synthetic held-out benchmark for executable workflow synthesis. RIPPLE improves validation success by up to 23.1% and yields positive gains on two additional frozen language-model backbones, while maintaining edit efficiency and low execution cost. Targeted interaction analysis further demonstrates both properties: a segment-local tool-use edit changes downstream resource resolution and validation, while an edit beneficial in isolation becomes harmful after composition.

补充信息

↑