arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解并利用基于反馈的智能体规划中的初始化锚定弱点

Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning

Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Peng Zhan, Zheng Li, Shanqing Guo

arXiv 2609.29697首次发表:更新:

发表机构

Shandong University(山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示基于反馈的智能体规划存在初始化锚定弱点,提出黑盒攻击框架InitAnchor,利用方向偏移、上下文合理性和反证韧性信号,在112个任务中实现高达76.1%的攻击成功率,表明早期偏差难以被后续纠正消除。

AI 中文摘要

基于反馈的规划通过整合工具观察和纠正性反馈来提高智能体的可靠性。然而,其保护措施在规划各阶段可能并非均匀分布。我们对四种代表性反馈机制进行了逐轮分析,发现了一个初始化锚定弱点:第一轮反馈纠正了46%的对抗性方向,而在存活到接下来两轮的方向中,纠正率分别降至13%和7%。我们的分析将该弱点归因于三个相互作用因素:初始计划中上下文合理的偏移、反证不足,以及累积轨迹中已接受方向的持续性。基于这些发现,我们提出了InitAnchor,一个通过攻击者控制的外部材料利用该弱点的黑盒框架。它将这三个因素分别操作化为方向偏移、上下文合理性和反证韧性信号,在有限目标访问或无目标访问两种设置下均适用。在来自16个领域、6种智能体架构和5个骨干大语言模型的112个任务中,InitAnchor在两种设置下分别实现了76.1%和72.0%的平均攻击成功率(ASR),同时将第一轮缓解率分别降至21.0%和25.0%。它还对六种防御措施和六个真实世界智能体系统保持有效。这些发现表明,即使后续存在纠正,基于反馈的智能体仍可能保留早期偏差。

英文摘要

Feedback-based planning improves agent reliability by incorporating tool observations and corrective feedback. However, its protection may not be distributed uniformly across planning stages. We conduct a round-wise analysis of four representative feedback mechanisms and uncover an initialization anchoring weakness: the first feedback round corrects 46\% of adversarial directions, whereas the rates fall to 13\% and 7\% among directions surviving into the next two rounds. Our analysis attributes this weakness to three interacting factors: a contextually plausible shift in the initial plan, insufficient counterevidence, and the persistence of accepted directions in the accumulated trajectory. Based on these findings, we propose \textsc{InitAnchor}, a black-box framework for exploiting this weakness through attacker-controlled external materials. It operationalizes the three factors as directional-shift, contextual-plausibility, and counterevidence-resilience signals under either limited target access or no target access. Across 112 tasks from 16 domains, six agent architectures, and five backbone LLMs, \textsc{InitAnchor} achieves average ASRs of 76.1\% and 72.0\% under the two settings while reducing first-round mitigation rates to 21.0\% and 25.0\%, respectively. It also remains effective against six defenses and across six real-world agent systems. These findings show that feedback-based agents can retain early biases even when later correction is available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑