arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08303cs.AIcs.MA

通过轨迹投毒对自进化技能实施仅查询式后门攻击

Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning

Yuyang Luo, Haoran Wang, Kai Shu

AI总结:

本研究针对自进化技能系统提出仅查询式轨迹后门攻击TBA,通过构造查询诱导生成带后门技能,实验证实其可可靠植入条件后门且效果优于直接注入。

AI中文摘要:

智能体技能通过为复杂任务编码可复用流程来提升大语言模型(LLM)智能体的性能。然而,人工编写的技能往往难以适配长周期任务与不断变化的环境。为解决这一局限,研究人员开发出自进化技能系统,该系统可通过执行轨迹自动构建和更新技能,将技能获取从外部市场转移至可信的进化流程。通过用可信的内部构建替代外部技能获取,自进化技能系统降低了对依赖直接技能操纵的技能注入攻击的暴露风险。不过,这一技能进化流程可能引入新的攻击面,攻击者可通过智能体交互诱导受损轨迹,间接操纵技能进化。为验证该威胁,我们提出轨迹后门攻击(Trajectory Backdoor Attack, TBA),这是一种仅查询式攻击,可引导可信的技能进化流程生成带后门的技能。具体而言,我们精心构造攻击者提交的查询,引导智能体执行目标动作,并在轨迹中明确说明对应的激活条件。我们在多种触发任务中重复相同的条件-动作模式,同时保持干净查询不变,促使进化器将该模式整合为依赖可复用触发条件的规则,纳入进化后的技能中。在两个技能进化系统的三个基准测试上,使用四个开源和闭源骨干模型开展的实验表明,TBA可在保留干净任务效用的同时可靠植入条件后门,效果与直接技能注入相当甚至更优。该结果揭示了轨迹驱动的技能进化存在的关键漏洞。

英文摘要:

Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks. However, manually authored skills often adapt poorly to long-horizon tasks and changing environments. To address the limitation, self-evolving skill systems have been developed to automatically construct and update skills from execution trajectories, shifting skill acquisition from external marketplaces to a trusted evolution pipeline. By replacing external skill acquisition with trusted internal construction, self-evolving skill systems reduce exposure to skill injection attacks that rely on direct skill manipulation. However, this skill evolution pipeline may introduce a new attack surface in which an attacker can indirectly steer skill evolution by inducing compromised trajectories through agent interactions. To demonstrate the threat, we propose Trajectory Backdoor Attack (TBA), a query-only attack that steers a trusted skill-evolution pipeline toward producing a backdoored skill. Specifically, we craft attacker-submitted queries to lead the agent to perform the target action and explicitly state the corresponding activation condition in the trajectory. We repeat the same condition-action pattern across diverse triggered tasks, while leaving clean queries unchanged, encouraging the evolver to consolidate the pattern as a reusable trigger-dependent rule into the evolved skill. Experiments on three benchmarks across two skill-evolution systems using four open- and closed-source backbone models demonstrate that TBA reliably implants conditional backdoors while preserving clean-task utility, matching or even surpassing direct skill injection. The results reveal a critical vulnerability in trajectory-driven skill evolution.

↑